The evidence behind the screen

Research-built race intelligence, made usable before post.

Historical research supplies the weights, filters, and guardrails. The current card supplies today's runners and prices. GAMELIN keeps those roles separate so a large research corpus is never mistaken for live information about one race.

5,758,545 Starter-level records analyzed across multiple datasets spanning 1990–2026. This is not a count of unique horses and is not the row count loaded into a live read.

Four layers, each answering a different question.

Scale shows how broadly the ideas were tested. Validation shows what held up. Live inputs determine how strong today's read can be.

5.7M+Research corpusHistorical starter-level records used to study market behavior, pace, ratings, risk, and filtering.
318,702US validationEquibase starters across 42,618 races and 126 US tracks during 2023.
6,559Model validationRaces in the validated Conviction Index sample, including chronological holdout checks.
TodayLive race inputThe actual entry card, current prices, and any legal pre-race figures available for this race.

What was analyzed, where, and why.

Counts are starter-level rows unless explicitly labeled as snapshots. Licensed and raw datasets are audited locally but are not republished here.

DatasetPeriodGeographyCountPurpose
Archive3 historical results1990–2020UK, Ireland, Europe, and US records4,107,315 starters
395,186 races
Long-horizon market, rating, pace, field-size, and longshot research
Equibase US canon2023United States318,702 starters
42,618 races
126 tracks
US validation, calibration, and leakage-safe feature checks
Hong Kong / Singapore researchMulti-yearHong Kong and Singapore79,447 starters
6,348 races
Prior running style, projected pace, and international comparison
Racingformbook Duo2016–2026 partialUnited Kingdom and Ireland1,253,081 startersIndependent ratings-versus-market method validation
Timestamped live oddsMulti-yearHong Kong and Singapore260,042 snapshots
1,519 races
Price movement, drift, market-floor, and timing research; counted separately from starters

Corpus total: 4,107,315 + 318,702 + 79,447 + 1,253,081 = 5,758,545 starter-level records. The 260,042 odds observations are a separate snapshot dataset and are not added to that total.

The engine is built as much from failed ideas as winning ones.

The strongest product decisions came from removing signals that looked impressive but did not survive chronological, market-relative, or leakage-safe checks.

01
Market efficiency is the baselineUS win pools are difficult. The screen starts from normalized market probability rather than pretending the public price is irrelevant.
02
Longshots need suppressionVery low-probability runners repeatedly underperformed after price. GAMELIN treats longshot excitement as risk, not evidence.
03
Ratings matter relative to priceA strong rating is most useful when the market has not already overpaid for an obvious standout.
04
Pace must be projectedActual same-race running position is forbidden. Legal pace work must come from prior running style and pre-race information.
05
Live movement adds contextTimestamped prices help distinguish support from drift, but movement is never treated as a lock by itself.
06
Conviction gates actionThe pick names the most likely winner; the Conviction Index decides whether the evidence is strong enough for a bet, a smaller structure, or a pass.
Independent method proof, not a Del Mar profit claim.

The 1,253,081-runner UK/Ireland ratings study found 11/11 positive annual slices for a leakage-safe ratings-versus-market rule. The 2023+ test contained 14,993 selections and the tighter midprice subset contained 6,597. That validates the thesis that a clean pre-race rating can add value relative to price. It does not prove the same return in US pools, does not establish a Del Mar earnings claim, and is not used as a future guarantee.

Every report says what it actually used.

The historical corpus never replaces today's card. Input quality controls the strength of the current output.

Entry Card

Official entries + morning lines

Names the likely contenders and creates a watchlist. No current tote and no complete PP figures means confidence stays capped.

Live Odds

Current prices + line movement

Adds a stronger market read and true price discipline. Conviction remains limited if class and recent speed figures are missing.

Full Figures

Prices + class + recent speed

Activates the full v7 Conviction Index. The result is still probabilistic and can still recommend no wager.

Only information available before the race can influence a live read.

Backtests become fiction when post-race facts leak into the features. GAMELIN explicitly forbids them.

Forbidden post-race fields

  • Official finish position and payoffs
  • Current-race speed ratings
  • Actual first, second, stretch, or final call positions
  • Same-race RPR, Topspeed, and result comments
  • Any feature timestamped after the wager decision

Honest production limits

  • Entry cards usually lack a live tote and full past performances
  • Morning-line reads are watchlists, not strong value calls
  • International findings do not automatically transfer to US pools
  • Public PP availability varies by race and provider
  • Historical ROI is a portfolio-subset result, not a single-race promise

What the v7 Conviction Index validation actually says.

Across 6,559 validated races, the historical top pick won 30.8% of the time. The top-50% conviction subset won 35.5% and returned +10.6% in the historical portfolio backtest. A strict date holdout returned +5.7% at a 31.1% win rate while the market baseline returned -3.9%.

The scope of the claim matters.

The +10.6% figure describes repeated historical selections in the top half of the conviction distribution. It is not a return forecast, a guarantee, a claim that every top pick should be bet, or evidence that one current race will be profitable. GAMELIN can still say PICK ONLY, NO VALUE, VERIFY FIELD, WAIT, or PASS.

Auditable summaries, protected raw data.

The canonical ledger is reconciled against local run summaries, licensed Duo archives, the active model card, and the v7 validation report. Public methodology describes dataset provenance, row counts, tests, and constraints without republishing licensed race records or proprietary raw files.

Evidence manifest version: 2026-09-01. The 2026 corpus is partial through May 31, 2026. No guarantees. Information only.