← Back to blog

How We Predict NBA Games with Machine Learning

By ZenHodl. Dataset documentation, research and model evaluations are linked in the article. A separate filtered ledger of bot-attributed trades, including losses and its admission rules, is public at /results.

Our NBA prediction model processes live game states and outputs a calibrated win probability — updated every score change, every timeout, every substitution. Here's exactly how it works under the hood.

The Architecture

The prediction pipeline has four layers, each correcting the one above it:

  1. Base XGBoost model — trained on 40,000+ NBA game snapshots (2021-2026)
  2. Team stats overlay — live offensive/defensive ratings (ORtg, DRtg, pace)
  3. Injury overlay — 58 tracked star players via ESPN's injury API
  4. Isotonic calibration — post-hoc probability correction (held-out ECE 0.053 for NBA)

When LeBron gets ruled out 30 minutes before tip-off, our system automatically adjusts the Lakers' win probability by his impact factor (8%) before the game even starts.

The Training Data

We train on every NBA game from 2021-22 through 2025-26 — approximately 5,285 games with ~300 snapshots each. Each snapshot captures:

The model is a Split-Phase XGBoost — it learns different patterns for early-game (where Elo dominance matters most) and late-game (where score differential dominates).

Why Calibration Matters More Than Accuracy

A model can be highly accurate (picks the winner 75% of the time) but terribly calibrated (when it says 70%, teams actually win 55%). For trading on prediction markets like Polymarket, calibration is everything — because you're not picking winners, you're pricing probabilities. We cover this principle in depth in why your sports betting model loses money without calibration.

Our NBA held-out calibration metric (Expected Calibration Error) is 0.053 — meaning when the model says 70%, teams win roughly 65–75% of the time. That's solid out-of-sample calibration (the near-zero figures you'll see quoted elsewhere are usually in-sample fits). We achieve it through:

  1. Post-hoc isotonic regression on a held-out 2024-25 season calibration set
  2. Live rolling recalibration that auto-corrects every 25 resolved predictions

If the model starts drifting (maybe a rule change shifts scoring patterns), the live recalibrator catches it within days and corrects automatically. No manual intervention needed.

Real-Time Injury Adjustments

We track 58 NBA star players through ESPN's injury API with a 10-minute cache. Each player has a pre-computed impact factor:

When a player is listed as OUT, the model subtracts their impact from the team's win probability. When they're QUESTIONABLE, the adjustment is halved. The total adjustment per team is capped at ±15% to prevent extreme swings.

The Results

Over 175 live moneyline bot trades:

Correction (2026-09-25): the numbers above are an earlier, smaller-sample figure and do not reflect what a later fleet-wide audit found. That audit measured the bought-side calibration gap — realized win rate minus the market's own implied probability — across every sport we trade on executed fills. For NBA it is −31 percentage points, the most negative of any sport in the fleet, meaning realized outcomes have run well below what our entry price implied. We no longer describe the model as "genuinely better than the Polymarket crowd" for NBA; the honest reading is that the market ask is already calibrated and our own model has not demonstrated a positive edge on executed NBA trades. See our live results for the current filtered record.

Try It Yourself

The full prediction API is available at zenhodl.net/docs with a 7-day free trial. You get:

Or build your own from scratch with our 6-module bot course — Module 3 covers the exact XGBoost training pipeline described here.

See our live trading results for the filtered bot-attributed record, including losses and the page's measured execution-evidence coverage.

Related reading

Get ZenHodl Weekly

Dataset releases, research notes, and public results, including corrections.

Research and dataset updates from ZenHodl.

Want the data behind this post?

Historical sports prediction-market datasets with measured coverage, documented schemas, and disclosed gaps.