Research-built race intelligence, made usable before post.
Historical research supplies the weights, filters, and guardrails. The current card supplies today's runners and prices. GAMELIN keeps those roles separate so a large research corpus is never mistaken for live information about one race.
Four layers, each answering a different question.
Scale shows how broadly the ideas were tested. Validation shows what held up. Live inputs determine how strong today's read can be.
What was analyzed, where, and why.
Counts are starter-level rows unless explicitly labeled as snapshots. Licensed and raw datasets are audited locally but are not republished here.
| Dataset | Period | Geography | Count | Purpose |
|---|---|---|---|---|
| Archive3 historical results | 1990–2020 | UK, Ireland, Europe, and US records | 4,107,315 starters 395,186 races | Long-horizon market, rating, pace, field-size, and longshot research |
| Equibase US canon | 2023 | United States | 318,702 starters 42,618 races 126 tracks | US validation, calibration, and leakage-safe feature checks |
| Hong Kong / Singapore research | Multi-year | Hong Kong and Singapore | 79,447 starters 6,348 races | Prior running style, projected pace, and international comparison |
| Racingformbook Duo | 2016–2026 partial | United Kingdom and Ireland | 1,253,081 starters | Independent ratings-versus-market method validation |
| Timestamped live odds | Multi-year | Hong Kong and Singapore | 260,042 snapshots 1,519 races | Price movement, drift, market-floor, and timing research; counted separately from starters |
Corpus total: 4,107,315 + 318,702 + 79,447 + 1,253,081 = 5,758,545 starter-level records. The 260,042 odds observations are a separate snapshot dataset and are not added to that total.
The engine is built as much from failed ideas as winning ones.
The strongest product decisions came from removing signals that looked impressive but did not survive chronological, market-relative, or leakage-safe checks.
The 1,253,081-runner UK/Ireland ratings study found 11/11 positive annual slices for a leakage-safe ratings-versus-market rule. The 2023+ test contained 14,993 selections and the tighter midprice subset contained 6,597. That validates the thesis that a clean pre-race rating can add value relative to price. It does not prove the same return in US pools, does not establish a Del Mar earnings claim, and is not used as a future guarantee.
Every report says what it actually used.
The historical corpus never replaces today's card. Input quality controls the strength of the current output.
Official entries + morning lines
Names the likely contenders and creates a watchlist. No current tote and no complete PP figures means confidence stays capped.
Current prices + line movement
Adds a stronger market read and true price discipline. Conviction remains limited if class and recent speed figures are missing.
Prices + class + recent speed
Activates the full v7 Conviction Index. The result is still probabilistic and can still recommend no wager.
Only information available before the race can influence a live read.
Backtests become fiction when post-race facts leak into the features. GAMELIN explicitly forbids them.
Forbidden post-race fields
- Official finish position and payoffs
- Current-race speed ratings
- Actual first, second, stretch, or final call positions
- Same-race RPR, Topspeed, and result comments
- Any feature timestamped after the wager decision
Honest production limits
- Entry cards usually lack a live tote and full past performances
- Morning-line reads are watchlists, not strong value calls
- International findings do not automatically transfer to US pools
- Public PP availability varies by race and provider
- Historical ROI is a portfolio-subset result, not a single-race promise
What the v7 Conviction Index validation actually says.
Across 6,559 validated races, the historical top pick won 30.8% of the time. The top-50% conviction subset won 35.5% and returned +10.6% in the historical portfolio backtest. A strict date holdout returned +5.7% at a 31.1% win rate while the market baseline returned -3.9%.
The +10.6% figure describes repeated historical selections in the top half of the conviction distribution. It is not a return forecast, a guarantee, a claim that every top pick should be bet, or evidence that one current race will be profitable. GAMELIN can still say PICK ONLY, NO VALUE, VERIFY FIELD, WAIT, or PASS.
Auditable summaries, protected raw data.
The canonical ledger is reconciled against local run summaries, licensed Duo archives, the active model card, and the v7 validation report. Public methodology describes dataset provenance, row counts, tests, and constraints without republishing licensed race records or proprietary raw files.
Evidence manifest version: 2026-09-01. The 2026 corpus is partial through May 31, 2026. No guarantees. Information only.