Methodology
CLV evidence refreshed 2026-10-10

How we build win probability models

Game-state features, model versions and evidence you can inspect.

Model coverage: NBA, WNBA, NCAAMB, NCAAWB, CFB, NFL, NHL, MLB. Not every covered sport places live orders at any given time — several run in shadow evaluation while they earn (or re-earn) live status; the changelog tracks pauses. Research and course extensions are called out separately below.

Dated exported evidence

Read each exported cohort with its generation date, prediction basis, sample size and exclusions. A dated calibration report is not a current all-sport guarantee. The former CLV-gap claim was retracted July 21, 2026 and amended August 6.

Methodology explains the model design. Validation shows the current exported backtest snapshot, results show filtered public-ledger outcomes, and /clv publishes descriptive per-sport closing-line observations. For in-play entries that terminal measure can encode the outcome, so it is not presented as a model-skill or profitability gate.

Snapshot 2026-07-22 80 days old

Retraction · 2026-07-21

The "78-point CLV gap" claim has been retracted

This section used to present a 78.8-percentage-point CLV-conditioned win-rate gap (89.9% vs 11.2%) as empirical evidence of model skill. It was a measurement artifact: about 90% of our entries are in-play, and the captured close on those markets sits near the market's final resolution — so the split largely restates who won. The August 6 amendment explains that the opposite-side split is an algebraic identity, not an independently measured control. The flat-50c, shuffled-entry and coin-flip placebos are the empirical checks. The live, unconditioned aggregate and its exchange-price correction history live on /clv.

Read the full retraction →

Three views — what does each page show?

  • /validation — exported backtest snapshot. What the model would have done on a defined sample with current filters baked in.
  • /results — filtered public ledger. Confirmed positions under the published loader contract, including losses and nonbinary outcomes; the page also reports measured execution-identifier and direct-transaction-hash coverage.
  • /clv-evidence — retraction. Why the "78-pp CLV gap" we used to publish was a measurement artifact, with the August 6 distinction between the inverse-side identity and measured placebos.

ZenHodl Weekly

Follow dataset releases, research and corrections.

Updates on dataset releases, research notes, corrections and product changes.

For builders, traders and researchers following the data and published evidence.

Model estimates and market-derived references

A market-derived fair-value reference can be built by devigging sportsbook lines — averaging Pinnacle, FanDuel, and DraftKings odds. This gives you a consensus probability that tracks the market by construction. It can be useful for line shopping, but it is not an independent model view — because the output already agrees with the market.

The released game-state architectures described below use features such as score differential, clock, period, ratings and sport-specific inputs; some include a pregame probability prior. The API contract labels the exposed edge model_pre_blend: it is taken before the bot's market-blend step. That label does not establish that every upstream feature in every current model version is independent of market information. Inspect the named model, feature set and prediction basis in each dated evaluation before making an independence claim.

Model quality must be compared on matched games, prediction times and model versions. The dated public benchmarks show selected evaluations; their scores do not establish current all-sport performance. A separately evaluated model can provide another reference for comparison with market prices; its input independence and trading value must be assessed for that model and cohort. See /clv-evidence for the audit and retraction of our former outcome-conditioned CLV claim, /validation for the exported backtest snapshot, and /results for filtered public-ledger outcomes and evidence coverage.

Market-Derived ZenHodl model estimate
Inputs Sportsbook odds/lines Game state, ratings, sport features and pregame priors
Output Tracks market by construction API exposes the documented pre-blend estimate
Brier Score Measure on the same outcome sample See the dated paired benchmarks
Trading Value Depends on venue, price and execution costs Inspect filtered outcomes and descriptive CLV; neither is a model-skill or profit guarantee

Data pipeline

We scrape ESPN's play-by-play API across the core sports we model. Recorded states contain score, period, clock and other source fields. Row counts depend on source availability and the collection run; missing intervals remain missing.

Parquet
Storage format
Game state
One recorded row
Versioned
Source schema
Versioned
Evaluation cohorts

Sports: NBA, NCAAMB, NCAAWB, CFB, NFL, NHL, MLB

Data is stored as Apache Parquet files. One row = one game state (score, period, clock, ESPN WP, outcome label).

Module 1 of our course teaches historical collection for NBA, NCAAMB, NHL, NFL and CFB; production collectors and supported API sports can differ.

Feature engineering

The documented feature families below vary by sport and model version. We deliberately keep the feature set small — overfitting to noise destroys trading value.

Feature Source schema Description
score_diff All home_score − away_score
seconds_remaining All Total game seconds left
period All Current period/half/inning
time_fraction All Fraction of game elapsed (0→1)
elo_diff All Home Elo − Away Elo
pregame_wp All ESPN pre-game win probability (fixed prior)
score_diff_x_tf All Lead × time elapsed (interaction)
score_diff_sq All Lead² (quadratic, captures blowouts)
is_home_batting MLB 1 if home team is batting
down, distance CFB/NFL Football situation
yard_line CFB/NFL Field position
possession_home CFB/NFL 1 if home has the ball
pace features NBA total_score, ortg_diff, drtg_diff

Model architecture

We use sport-specific models — no one-size-fits-all approach. The sections below distinguish the current live platform from adjacent research and course material.

Public API coverage and documented model design

The models cover NBA, WNBA, NCAAMB, NCAAWB, CFB, NFL, NHL, MLB with sport-specific pipelines. Calibration and overlays depend on the model and serving path. Raw API model estimates and the bot's calibrated, market-blended trading values can differ; the API/bot disclosure describes that distinction.

Basketball & Football (NBA, NCAAMB, NCAAWB, CFB, NFL)

Split-Phase XGBoost with 16 features including team offensive/defensive ratings (ORtg, DRtg), pace, momentum (scoring runs over last 2 and 5 minutes), Elo ratings, and interaction terms (score_diff × time_fraction). Evaluate each released model with its own training dates, chronological splits and held-out cohort; static architecture descriptions are not a current quality report.

Injury overlay: The documented NBA overlay uses ESPN injury status with a 10-minute cache and configured player impact weights. These weights are modeling heuristics, not measured causal injury effects or proof that the upstream status is fresh. When a player is OUT, the model subtracts their impact; QUESTIONABLE halves the adjustment. Total cap: ±15%.

Model selection uses a hybrid criterion: near-best Brier score AND best trading value (c/trade) on a held-out backtest window. This is a selection criterion; it does not guarantee better live P&L or prevent overfitting to the backtest.

Hockey (NHL)

XGBoost with 17 features including all basketball features plus hockey-specific metrics: power play %, penalty kill %, save %, faceoff win %, and penalty minutes differential. Inspect the corresponding model report for its dated test sample and baseline; these features do not establish a current performance level.

Injury overlay: The documented ESPN skater-status overlay excludes goalies, which use a separate adjustment. Player weights are configured heuristics; status availability and freshness depend on the source.

Goalie & shot overlay: Starting goalie detection adjusts for goalie quality. Expected goals (xG), Corsi, and power play state provide real-time shot quality signals. Combined cap: ±12%.

Baseball (MLB)

Ensemble model with starting pitcher ERA, WHIP, K/9 as base features. A model report must state the test dates and sample size for any Brier or ECE result. Three post-model overlays stack on top of the base prediction:

Soccer (Bundesliga, EPL, La Liga, Serie A, Ligue 1)

The published April 2026 soccer incident describes a live-trading freeze (force-shadowing). The XGBoost model trained on 2,949 matches had silently vanished from the server in April 2026, degrading the live bot to a Poisson fallback that measures worse than its own entry prices. A retrained challenger failed its promotion gates, and the response included artifact-presence monitoring. Current status belongs in the repair scorecard and changelog. Pregame research continues on the Poisson + XGBoost stack.

Esports (CS2, LoL)

Both esports run models on the platform; live order flow varies with each sport's current risk status. The published August 2026 record describes a CS2 sizing pause and a LoL revert to v1 after v2 lost two head-to-head evaluations (see below). Check the repair scorecard and changelog for subsequent status.

CS2: 15 features including round differential, Elo, map-specific win rates (per team per map), recent form, head-to-head record, and Elo momentum. Economy data from bo3.gg and HLTV scorebot provides equipment value and consecutive-loss tracking as post-model overlays. The documented CS2 entry configuration used underdogs (20–42c) and a 20c+ minimum modeled gap after a cohort audit. Those are configuration choices, not proof that future underdog entries will outperform.

LoL: v1 analytical model — Elo ratings plus a binomial series model for BO1/BO3/BO5. The v2 XGBoost override (20 features: gold differential, kills, towers, dragons, barons, inhibitor advantage, and more) was reverted in August 2026 after two independent evaluations on 110 fire-anchored matches both favored v1 (Brier 0.1692 vs 0.2077). That dated comparison supports the reported revert; it does not establish the status or performance of later versions.

Tennis (ATP, WTA)

Hierarchical analytical model: player-specific serve win rates (by surface) → point probability → game probability → set probability → match probability. Serve rates are computed from 13,174 ATP matches (2020–2024) using the Sackmann dataset, giving us player-specific rates for hard, clay, and grass surfaces instead of tour averages. The documented Elo ratings are surface-aware; player coverage depends on the rating-file version. The documented activation threshold is 50 resolved trades; reaching it does not establish adequate out-of-sample calibration.

Calibration, measurement and risk controls

The documented bot pipeline can apply calibration and overlays before its measurement and risk checks. Applicability varies by sport, venue and serving path; the raw API model output is not identical to this chain:

  1. Post-hoc isotonic calibration: Fitted on held-out calibration season. Evaluate the fitted map on an unseen cohort with stated bins and sample counts. A low fitting-set error does not guarantee held-out calibration, and a model estimate of 70% is not a promise of a 70% win rate in a new sample.
  2. Live rolling recalibration: Automatically refits an isotonic correction every 25 resolved trades per sport, using a 500-sample rolling buffer persisted to disk. This is a documented refresh mechanism, not proof that every drift is detected or corrected; each resulting map still needs an independent evaluation.
  3. Closing-line value (CLV) audit: Admitted live fills are measured against our captured closing price where one exists; the scorecard publishes the current measured count and coverage rather than freezing an old percentage into this methodology page. The winner/loser split we used to cite here was retracted 2026-07-21 as a measurement artifact; the honest unconditioned aggregate (share of fills beating the close and mean CLV per trade) is published per sport at /clv. The four-failure-mode taxonomy (bad model / bad blend / bad timing / bad CLV label) is published at /clv/repair alongside the explicit reactivation criteria for any sport currently paused.
  4. Dynamic position sizing: Rolling 50-trade performance per sport determines a size multiplier (0.25x–1.5x). Hot sports (>55% WR + positive P&L) get sized up; cold sports (<40% WR) get sized down. Recomputed every 10 minutes.
  5. Meta-model edge filter (research, shadow mode only): A Gradient Boosting classifier trained on resolved trades was proposed to separate real edges from noise. It has never been trusted to block live trades — it runs in shadow mode as a research instrument. The "filters 28% of losing trades" figure we once cited was not reproducible and has been removed.

Spread/Total

Spread and total models require a different setup from moneyline markets: regression on remaining margin or total, then a distributional layer to convert that forecast into cover/over probabilities.

Backtesting methodology

Our backtests are designed to avoid the common mistakes that inflate results.

What we tried that doesn't work

We believe in showing failures alongside successes.

Choose your next step

Read the proof, inspect live results, or go straight to the live platform.