← Back to blog

ESPN API in Python: Scoreboard, Summary and Win Probability

updated 2026-10-06 espn python tutorial historical-data win-probability

By ZenHodl. Dataset documentation, research and model evaluations are linked in the article. A separate filtered ledger of bot-attributed trades, including losses and its admission rules, is public at /results.

If you landed here after pasting site.api.espn.com/apis/site/v2/sports/basketball/nba/summary?event=EVENT_ID into a search box, this is the answer you were looking for.

We run this endpoint in production across eleven sports to drive live win-probability models, at roughly 120,000–190,000 requests a day from a single host. Everything below is measured, not inferred — verified against live responses on 13 August 2026 and against months of production logs. Where a statement is about our configuration rather than ESPN's behaviour, it says so.

The short answer

winprobability is a top-level array on the summary response, a sibling of header, plays and boxscore. Every element carries exactly three keys:

{
  "homeWinPercentage": 0.457,
  "tiePercentage": 0.0,
  "playId": "4018165000001990057"
}

Three things that trip people up immediately:

Which sports actually have it

This is the question the docs can't answer, because there are no docs. Verified by live request on 2026-08-13, one completed game per league:

Sport winprobability? Entries in sample
NBA Yes 457
WNBA Yes 417
NCAA men's basketball Yes 397
NCAA women's basketball Yes 529
NFL Yes 186
College football Yes 155
MLB Yes 77
NHL No — key absent entirely —
Soccer (EPL, MLS, UCL tested) No —
Tennis No —

Across NHL, EPL and MLS, 14 of 14 completed games sampled had no winprobability key at all — not an empty array, the key simply isn't there.

MLB has win probability. This is worth stating loudly because the opposite is widely believed — and we were part of the problem. Our own scraper hardcodes "has_wp": False for MLB and writes espn_home_wp: None on every row, which is why our internal reports label MLB as having no ESPN data. That was our bug, not ESPN's absence. Eight of eight completed MLB games sampled returned a populated array.

NHL is a genuine absence, and ESPN is explicit about it. The core API returns HTTP 400 with the message Probabilities are not supported for sport: hockey, league: nhl. Soccer returns the same shape of refusal.

Granularity differs by sport, and it matters

The pregame element is a special case

For football, winprobability[0] is a synthetic pregame entry whose playId is the game ID with 1 appended — it joins to no play. For basketball, element 0 is the opening tip (period 1, clock 12:00, 0–0). For MLB it is already the result of the first plate appearance — there is no pregame row.

For a game that hasn't started, the key is present but the array is empty. The pregame number lives in a separate top-level predictor block that only appears before the game starts.

Four traps that will corrupt your dataset

1. The last element is a label, not a forecast. On a completed game, winprobability[-1] is always exactly 1.0 or 0.0. If you train on the full array you are feeding the model the answer.

2. ESPN emits exact 0.0 and 1.0, not clipped epsilons. In our NBA corpus that's 455 rows at exactly 1.0 and 396 at exactly 0.0. Log-loss is infinite unless you clip.

3. Event IDs are not chronological. You cannot infer a date range by sorting them. In our NCAAMB corpus, event 401823757 is dated 2025-11-12 while the numerically lowest ID 401804830 is dated 2025-12-21.

4. Precision varies by sport. homeWinPercentage is quantised to 3 decimals for basketball and MLB (exactly 1,001 distinct values across hundreds of thousands of rows) but 4 decimals for NFL.

The User-Agent rule is backwards

This is the single most useful operational fact on this page, and it inverts the standard advice about scraping.

Send a stock HTTP-library User-Agent. Do not send a browser one.

Verified on two independent networks:

User-Agent Result
curl/8.5.0, python-requests/2.31.0, aiohttp/3.9.5 200
Go-http-client/1.1, okhttp/4.9.3, axios/1.6.0, Java/17.0.1 200
Mozilla/5.0, full desktop Chrome UA 403
Empty UA, PostmanRuntime/7.39.0, SomethingBot/1.0 403
A polite descriptive UA with a contact URL 403

That last row cost us. Our benchmark cron sends ZenHodlBenchmark/1.0 (+https://zenhodl.net/benchmarks) — the well-behaved, RFC-spirited thing to do — and has received 179,775 consecutive 403s. Only a complete browser header set (Accept, Accept-Language, Referer, Origin, sec-ch-ua, Sec-Fetch-*) gets a browser UA through; no single header rescues it. Our production score clients send no headers at all, and that is exactly why they work.

Two things make this failure mode nasty:

Rate limits, caching and cost

Request Observed max-age
Live scoreboard (today, games in progress) 1–10s
Summary of a completed game 4–7s
Scoreboard for a past, settled date ~52–54s

For live polling that single-digit window is the real ceiling on useful frequency — anything faster returns the same bytes. For backfilling settled dates you get about a minute of caching for free. - No ETag, no Last-Modified. Conditional requests are not available; the full body ships every poll. - Payloads are large but compress ~14×. An MLB scoreboard is 443,945 bytes decoded, 30,767 gzipped. Make sure your client actually sends Accept-Encoding. - Reading one win-probability number costs ~450KB. An NFL summary was 460,965 bytes for 188 entries plus 18 other top-level keys. Cache per game. - Latency is excellent from a datacenter — 13ms min / 15ms median over warm keepalive fetches, versus 137–261ms residential.

Two intermittent failure modes worth defending against: ESPN occasionally returns HTTP 200 with a blank Content-Type, which crashes aiohttp's strict JSON decode while curl and requests parse it fine (3 times in 14.5 days). And real-world flakiness presents as TCP connect timeouts, never HTTP errors — 45 connect failures and 16 timeouts over the same window, with zero HTTP status errors.

A scoreboard parameter nobody documents

For men's college basketball, groups=50 is what expands the scoreboard, not limit. The default response is truncated to a subset of the day's games; adding groups=50 (Division I) returns the rest. limit=300 and even limit=1000 return a response byte-identical to the default.

The exact counts depend on the date you query — on 2026-02-10 the default gave 11 events and groups=50 gave 22; on a busier date we saw 7 versus 62. What is stable is the behaviour: groups expands, limit does nothing.

How accurate is ESPN's win probability?

A code snippet can tell you the field's shape. It cannot tell you whether the number is any good. We measured it.

Method: take ESPN's homeWinPercentage at each play it publishes one for, against the binary realised outcome from the final score in the same payload. It asks one question — when ESPN said 70%, did the home team win 70% of the time? No betting market, no model of ours, involved.

In-game Expected Calibration Error, 10 equal-width bins:

Sport ECE Games Snapshots
NCAA men's basketball 0.51% 5,345 1,416,969
NCAA women's basketball 0.77% 5,344 2,474,806
College football 1.27% 946 166,059
WNBA 2.16% 312 120,895
NBA 3.10% 888 329,817
NFL 3.40% 285 40,970

The short verdict: ESPN's in-game win probability is well calibrated. If you need a reference probability for a live game and you are not trying to beat the market, it is good enough to build on.

Now the caveats, because the headline numbers flatter it:

Where it is measurably off

Two findings survive a bin-free calibration-in-the-large test with bootstrap intervals excluding zero:

One caveat we have to state on those band-level results: 38 non-empty bands were tested at the 95% level, so ~1.9 false positives are expected by chance and 6 were observed. Only the two large NCAAWB bands (+7.9pp on n=405, +6.6pp on n=442) have effect sizes that make chance implausible.

And the NFL pre-game figure of 9.38% ECE — the worst number in the set — rests on 285 games, with 19% of the error coming from bins holding fewer than 30 games, including bins of n=2 and n=5. Report it with its n or not at all.

These are frozen-corpus numbers, from single-season scrapes bounded by our own collection windows (NBA 2025-26, NCAAMB/NCAAWB 2025-26, NFL 2024, CFB 2024, WNBA 2025), not a live monitor.

ESPN win probability API: working NBA Python example

Install requests if needed (python -m pip install requests). This reads one completed NBA game, Indiana at Oklahoma City on the June 22, 2025 scoreboard. The public summary request was rechecked without authentication on October 6, 2026. For another game, get its event ID from the NBA scoreboard: https://site.api.espn.com/apis/site/v2/sports/basketball/nba/scoreboard?dates=YYYYMMDD. The verified date query used 20250622; an event ID is not a date.

Read winprobability[*].homeWinPercentage on this summary endpoint, rather than guessing a gameprob field. In the verified NBA response, it is a 0–1 fraction: 0.591 displays as 59.1%, not 0.591%. Other endpoints' field names do not establish their units. Join each playId to plays[*].id to get the observed period and game clock.

No custom headers. That is deliberate — see the User-Agent section.

import requests

URL = "https://site.api.espn.com/apis/site/v2/sports/basketball/nba/summary"
response = requests.get(URL, params={"event": "401766128"}, timeout=15)
response.raise_for_status()  # Check status before trying to decode JSON.
summary = response.json()

entries = summary.get("winprobability") or []
if not entries:
    raise ValueError("No win-probability observations in this response")
plays = {str(play["id"]): play for play in summary.get("plays", [])}

for entry in entries[:5]:
    home = float(entry["homeWinPercentage"])
    tie = float(entry["tiePercentage"])
    if not (0 <= home <= 1 and 0 <= tie <= 1 and home + tie <= 1):
        raise ValueError("Expected 0–1 fractions; inspect the response schema")
    away = 1 - home - tie  # Derived value, not a site-summary response field.
    play = plays.get(str(entry["playId"]))
    if play is None:
        continue  # Do not invent a clock for an unmatched probability.
    period = play.get("period", {}).get("number")
    clock = play.get("clock", {}).get("displayValue")
    print(entry["playId"], period, clock,
          f"home={home:.1%}", f"away={away:.1%}", f"tie={tie:.1%}")

This prints the first five observations, not five independent forecasts. For model evaluation, exclude the completed game's terminal outcome entry and define the information cutoff before joining any other data. Treat a missing or empty array as "no data" rather than substituting 0.5, and clip to [1e-6, 1-1e-6] before any log-loss. The observed NBA shape is not a guarantee for another sport or a future response.

The other ESPN API

There is a second, parallel API at sports.core.api.espn.com with a probabilities resource. It returns four fields the site summary drops — awayWinPercentage, sequenceNumber, lastModified and source — plus, for NBA only, spreadCoverProbHome, spreadPushProb and totalOverProb.

The trade-off is paging. The site summary ships all 457 NBA entries in one response; the core API paginates at 25 items and needs 92 requests for the same game. Use the site endpoint unless you specifically need awayWinPercentage or the spread probabilities.


Both APIs are undocumented and unofficial. Nothing here is a contract — ESPN changed its User-Agent handling abruptly at 15:05 UTC on 4 August 2026 with no notice, and we found out through a log. Wrap it in timeouts, schema checks and alerts that fire on 4xx.

Next: study probabilities alongside market prices

Calibration alone does not establish a tradable edge. If your next question is how market prices moved, inspect captured quotes separately and check the timestamps and coverage before choosing an alignment method.

ZenHodl's market-data downloads contain prediction-market observations, not ESPN winprobability history. Matching ESPN events and probabilities to those observations is a separate research step. The free sample, schema and small tryout below let you check whether the data fits that step before considering a larger archive.

We publish the calibration numbers because they are checkable. Our own trading results, including the losses, are public.

Optional next step: historical market data

Check the data before you commit

For research on Kalshi market prices, start with the fields and recorded coverage. These downloads contain sampled books and captured trade prints; ESPN probabilities are a separate input.

  1. 1. Inspect free rows

    Open a depth CSV; read the schema for book and trade fields and known gaps.

    Download free depth CSV Read the schema
  2. 2. Test a small download

    The $9 Kalshi tryout contains two fixed day partitions. Review its own coverage before buying.

    See the $9 tryout
  3. 3. Choose more coverage

    If the data fits your workflow, compare the full tape's recorded window and limitations.

    Review the archive

Related reading

Get ZenHodl Weekly

Dataset releases, research notes, and public results, including corrections.

Research and dataset updates from ZenHodl.