Working Paper · July 2026

Exploiting Short-Term Inefficiencies in Sports Prediction Markets

Seven-sport historical core trained on 41,000+ games. The 2,625-row result is an ESPN-WP proxy simulation, with corrections preserved in the record.

Version v1.2 Game-clustered CI includes zero Correction record

Version 1.2 — Corrected July 21, 2026. This revision corrects the provenance and interpretation of the 2,625-row historical evaluation, retracts the earlier independence-based significance claim, and retracts the outcome-contaminated final-close CLV winner/loser comparison. The correction record is part of the paper; the earlier PDF is retained only as a superseded notice.

Scope note (added 2026-09-27): this paper evaluates the 7 sports listed in §2.1. WNBA was added to the live engine later; nothing below covers it.

ZenHodl Research | July 2026 Version 1.2 (Corrected July 21, 2026)


Abstract

This paper describes a pipeline for producing calibrated sports win probabilities and using them in a prediction-market trading system. It documents data acquisition, model training, calibration, signal generation, execution, and the limitations found during later audits.

The historical core described here covers seven ESPN-based sports and 41,000+ games. It fits logistic-regression and gradient-boosted candidates, calibrates them on chronological calibration data, and evaluates win probabilities at multiple in-game states. Separate live-sport architectures are identified where they appear; they are not silently pooled into the historical artifact.

The published 2025-26 artifact contains 2,625 simulated opportunities and a 69.8% outcome win rate. It is an ESPN win-probability proxy simulation, not a Polymarket or Kalshi order-book backtest: the simulated entry is ESPN WP plus a 0.5-cent half-spread. The artifact reports +5.9 cents per opportunity relative to ESPN WP, +5.4 cents after that half-spread, and +2.4 cents after a further assumed 1.0 cent of slippage and 2.0 cents of taker cost. After clustering correlated rows by game, the 95% interval for mean simulated net P&L includes zero. These numbers therefore do not establish executable market alpha.

The March-April 2026 live table is retained as a dated historical snapshot. Current confirmed live results and the CLV-gap retraction are maintained separately at https://zenhodl.net/results and https://zenhodl.net/clv-evidence.


1. Introduction

1.1 Prediction Market Efficiency

Prediction markets aggregate information through trading to produce probability estimates for future events. Theory suggests these markets should be approximately efficient, with prices reflecting the true probability of outcomes (Wolfers & Zitzewitz, 2004; Arrow et al., 2008).

However, efficiency is not instantaneous. During live sporting events, new information arrives continuously through score changes, momentum shifts, and game clock progression. Market participants process this information at varying speeds, creating brief windows where prices lag the true state of the game.

1.2 The Information Latency Hypothesis

Our core hypothesis is that during live games, there exists a 15-60 second window after significant game events (score changes, period transitions, possession changes) where prediction market prices have not fully adjusted to the new game state. This latency arises from:

  1. Human processing delay: Most market participants watch games on television with inherent broadcast delay
  2. Attention fragmentation: Participants monitoring multiple games cannot react to all simultaneously
  3. Asymmetric information integration: Score changes are immediately observable, but their probabilistic implications require computation

A machine learning model that processes game state features in real-time can compute updated win probabilities faster than the median market participant, capturing the information premium during this adjustment window.

Foundational work on sports probability and market efficiency informs our approach. Stern (1991) established the statistical framework for modeling win probability as a function of in-game state variables, providing the foundation that underlies modern win probability models including ours. Sauer (1998) surveyed the economics of wagering markets comprehensively, documenting both the surprising efficiency of traditional betting markets and the specific conditions under which inefficiencies persist. Wolfers and Zitzewitz (2004) provided an influential overview of prediction markets and their information aggregation properties.

Empirical studies have documented specific inefficiencies exploitable by quantitative approaches. Borghesi (2007) demonstrated persistent biases in NFL betting markets related to home-field advantage and weather effects, providing evidence that even mature sports betting markets are not fully efficient. Croxson and Reade (2014) examined in-play betting market efficiency around goal arrivals in soccer, finding rapid but not instantaneous price adjustment. Kaunitz, Zhong, and Kreiner (2017) showed that systematic exploitation of closing line value in traditional sportsbooks can yield positive returns, though bookmaker countermeasures limit scalability.

Most relevant to our work, Page (2012) documented biases in prediction markets during live events. Our system applies calibrated sports models to markets that settle on Polygon through Polymarket. Whether any information-latency effect remains executable after quote freshness, fill selection, fees, and game-state changes is an empirical question; the historical proxy artifact does not answer it.


2. Data and Methodology

2.1 Data Sources

Game State Data: We polled ESPN's public API every 5 seconds for all live games across 7 sports: NBA, NFL, NHL, MLB, NCAAMB (men's college basketball), NCAAWB (women's college basketball), and CFB (college football). For each game, we extract:

  • Score differential
  • Period/quarter/inning
  • Time remaining (seconds)
  • Possession (football only)
  • Down, distance, yard line (football only)
  • Starting pitcher ERA, WHIP, K/9 (baseball only)
  • Power play/penalty kill status (hockey only)

Elo Ratings: We maintain continuously-updated Elo ratings (Glickman, 1999) for all teams using a K-factor of 20, home-court advantage of 50 points, and 50% seasonal regression. Ratings are computed from historical results and updated after each completed game.

Market Prices: Real-time bid/ask prices from Polymarket's WebSocket feed, with additional venue coverage from Kalshi and OddsAPI (DraftKings, FanDuel, BetMGM) for multi-venue comparison.

Historical Evaluation Proxy: The 2,625-row artifact discussed in §5.1 contains no exchange order-book prices. Its comparison price is ESPN's in-game win probability, transformed into a simulated ask by adding 0.5 cents. Production market feeds and the historical proxy evaluation are distinct datasets and must not be treated as interchangeable.

Training Data: 41,000+ historical games across all 7 sports, spanning the 2020-21 through 2025-26 seasons. Each game produces multiple evaluation snapshots at different game states, yielding hundreds of thousands of training examples. Each snapshot contains game state features paired with the actual binary outcome (home team win/loss).

2.2 Model Architecture

For each sport, we train an ensemble of two model classes:

  1. Logistic Regression with Natural Spline Features: Provides a well-calibrated baseline with interpretable coefficients. Spline transformations on score_diff and time_fraction capture non-linear relationships (e.g., a 10-point lead means different things in the first quarter versus the fourth).

  2. Gradient-Boosted Trees (XGBoost): Captures complex feature interactions. Trained with max_depth=4, learning_rate=0.1, n_estimators=200, and regularization (lambda=1.0, alpha=0.1) to prevent overfitting.

Both models are post-hoc calibrated using isotonic regression on a held-out calibration set. The final ensemble weights are determined by minimizing Brier score on the calibration set.

We assessed model sensitivity to key hyperparameters. Brier scores are stable within +/-0.005 across max_depth in {3, 4, 5} and learning_rate in {0.05, 0.1, 0.2} for all sports. The ensemble weights between logistic regression and XGBoost were determined by minimizing Brier score on the calibration set, with typical weights of 40-60% XGBoost depending on the sport.

Feature Engineering: - score_diff: Home score minus away score - time_fraction: Fraction of game remaining (1.0 = start, 0.0 = end) - score_diff_x_tf: Interaction term capturing how score leads change in importance over time - score_diff_sq: Squared score differential for non-linear response - elo_diff: Pre-game Elo rating difference - Sport-specific features as described in Section 2.1

2.3 Temporal Split Methodology

We use temporal splits at the season level. For sports with 3+ seasons of data, the oldest seasons are used for training, the second-newest season for calibration, and the most recent season for testing. For sports with fewer seasons, we use a 60/20/20 chronological split within the available data. Historical artifacts compared candidate families on the outer test season, so their selected-model metrics remain mildly optimistic. On July 21, 2026, the trainer default was changed to a precommitted calibration-fitted ensemble; legacy outer-test selectors now require an explicit research-only override. Nested rolling-origin selection with a final untouched season remains the stronger standard for a clean performance claim.

Elo ratings are computed in a walk-forward manner, using only games completed before the current evaluation point.

2.4 Calibration

Calibration is critical for our application. A model that is discriminative but poorly calibrated will overestimate or underestimate true probabilities, leading to systematic trading errors.

We use isotonic regression calibration (Zadrozny & Elkan, 2002) because it makes no parametric assumptions about the calibration function. We measure calibration quality using:

  • Expected Calibration Error (ECE): Weighted average of absolute calibration error across probability bins
  • Brier Score: Proper scoring rule that measures both discrimination and calibration
  • Reliability Diagrams: Visual assessment of calibration across the probability range

2.5 Uncertainty Quantification

Each model includes an uncertainty estimate based on calibration-error-based uncertainty bands. For each time-fraction bucket, we compute the average absolute calibration error on the held-out calibration set. This provides an empirical estimate of model uncertainty that varies by game state — wider early in games when outcomes are less determined, and narrower late in games with large score differentials.

For each prediction, we provide:

  • A point estimate of win probability
  • A confidence interval width that varies by game state
  • Early-game predictions have wider intervals (more uncertainty)
  • Late-game predictions with large score differentials have narrower intervals

This uncertainty estimate informs position sizing: we trade smaller when uncertainty is high and larger when the model is confident.


3. Model Performance

3.1 Overall Metrics

Sport Brier Score ROC-AUC ECE Training Games
NCAAWB 0.111 0.910 0.026 11,581
CFB 0.121 0.905 0.015 2,411
NBA 0.128 0.900 0.053 5,285
NCAAMB 0.146 0.865 0.024 12,285
MLB 0.152 0.859 0.014 4,413
NFL 0.174 0.817 0.092 1,140
NHL 0.205 0.739 0.034 4,225

NCAAWB achieves the lowest Brier score (best overall accuracy), while NHL has the highest (hardest to predict due to the low-scoring, high-variance nature of hockey). NFL shows the highest ECE (0.092), indicating the most room for calibration improvement; its small training sample (1,140 games) and high week-to-week variance make calibration harder than for the high-volume basketball and baseball models.

3.2 Calibration Analysis

Most rows in the historical table have aggregate ECE at or below ~0.03; NBA (0.053) and NFL (0.092) are higher. ECE is a bin-weighted summary, not a guarantee that every quoted probability band is calibrated, and these selected-model metrics inherit the model-selection caveat in §2.3.

3.3 Uncertainty Tables

Each model includes a lookup table mapping game-state time fractions to expected uncertainty widths. For example, in NBA:

  • Early game (75%+ remaining): Uncertainty width 0.081 (high)
  • Mid game (25-75% remaining): Width 0.040-0.060
  • Late game (<25% remaining): Width 0.024 (low)

These widths inform the confidence level assigned to each trade signal.


4. Edge Detection and Execution

4.1 Signal Generation

For each live game with a matched Polymarket market, we compute:

edge_c = fair_wp_c - market_ask_c

Where fair_wp_c is the model's fair win probability in cents (0-100) and market_ask_c is the current Polymarket ask price.

A trade signal is generated only after sport- and mode-specific gates. Representative gates include: - edge_c >= min_edge (configured by sport and evidence state; not one fleet-wide floor) - fair_wp_c is between 55 and 95 cents (avoid extreme probabilities) - Market spread is less than 6 cents (liquidity filter) - The process-local tracker copy passes its existing age gate - The model's uncertainty width is below a sport-specific threshold

4.2 Execution

Trades are submitted as Fill-and-Kill (FAK) marketable limit orders through Polymarket's CLOB API; partial fills are accepted and the unfilled remainder is cancelled, while the market itself settles on Polygon. (Corrected 2026-07-21: earlier versions said Fill-or-Kill; production has posted FAK, and partial fills are first-class ledger facts.) Representative execution parameters include:

  • Slippage tolerance: 2 cents
  • Maximum entry price: 78 cents
  • Position sizing: Kelly criterion at quarter-Kelly with maximum bet caps
  • Concurrent position limit: 8 positions maximum

The tracker-copy age above is not a complete frozen-quote defense: polling can re-stamp an unchanged cached quote. As of this correction, final-signal and pre-submit paths record a separate local WebSocket receipt age and a would_block label, but do not enforce it. It is not an exchange event timestamp. Enforcement requires a pre-specified, freshness-qualified shadow cohort showing that the rule improves outcome-blind markouts without rejecting legitimately quiet books.

4.3 Execution Cost Model

The historical proxy simulation applies the following assumed friction relative to ESPN WP: - Simulated half-spread: 0.5 cents, included in each artifact entry_c - Assumed slippage: 1.0 cent, deducted from gross P&L - Assumed taker cost: 2.0 cents, deducted from gross P&L - Total assumed friction relative to ESPN WP: 3.5 cents per opportunity

The artifact does not contain a market quote, measured latency, depth, queue position, partial fills, fill failures, or realized fees. These assumptions are arithmetic stress adjustments, not measured execution costs.

4.4 Multi-Venue Comparison

Production services monitor several market and sportsbook sources for comparison and research. The 2,625-row artifact does not evaluate multi-venue routing, executable cross-venue prices, or routing performance. No optimal-routing conclusion is drawn from it.


5. Results

5.1 ESPN WP Proxy Simulation (2025-26 Season)

Metric Value
Simulated opportunities 2,625
Win rate 69.8%
Gross relative to ESPN WP +5.9c/opportunity
Simulated half-spread, included in entry_c -0.5c
Gross after simulated half-spread +5.4c/opportunity
Assumed slippage -1.0c
Assumed taker cost -2.0c
Simulated net after all assumed friction +2.4c/opportunity
Total net P&L +$62.69 (computed trade-by-trade; the 2.4c average is rounded)

Provenance correction (2026-07-21). A row-by-row audit of the published artifact found that every entry price is the matching ESPN side win probability plus the configured 0.5-cent half-spread, within stored rounding. The artifact's method matches api/backtest_engine.py: ESPN WP is the comparison proxy, not an exchange order book. Earlier descriptions of these rows as Polymarket prices, Kalshi prices, real market snapshots, or filled trades are retracted.

This simulation can compare the model's forecasts with ESPN's forecasts under an assumed cost schedule. It cannot measure fillability, adverse selection, market impact, venue fees, or realized trading profit.

By Sport:

Sport Opportunities Win Rate Gross c/Opportunity after half-spread
NCAAMB 1,237 76.6% +9.3c
NCAAWB 864 66.2% +2.2c
NFL 286 58.0% -1.0c
NBA 238 61.3% +3.9c

The aggregate +2.4-cent figure includes all 3.5 cents of assumed friction relative to ESPN WP. Per-sport figures above are descriptive gross values after the simulated half-spread and before the additional 3.0-cent deduction.

NHL, MLB, and CFB produced zero qualifying proxy disagreements in this artifact. Because it contains no exchange book, that result says nothing about Polymarket or Kalshi coverage or liquidity.[^1]

NFL is the only sport with negative expected value in the backtest, likely due to a smaller training sample (1,140 games) and the NFL model's previously identified temporal split issue (since corrected).

NCAAMB accounts for 47% of the simulated opportunities. That concentration limits even the forecast-proxy interpretation; the artifact does not demonstrate broad market profitability.

[^1]: The 4 sports shown (NCAAMB, NCAAWB, NFL, NBA) account for all 2,625 trades. NHL, MLB, and CFB had zero qualifying trades during this period.

5.1.1 Statistical Significance

Correction (2026-07-21). Earlier versions treated 2,625 rows as independent and claimed significance at the 1% level. The rows cover only 1,397 unique games (~1.9 opportunities per game), so same-game observations must be clustered. Using the artifact's own simulated net P&L:

t = 2.39 / 1.34 = 1.78 (game-clustered, 1,397 clusters; two-sided p ≈ 0.075) 95% CI for mean net P&L per trade: [−0.24c, +5.01c]

The interval includes zero. The simulated edge is not statistically distinguishable from zero at conventional levels, and even this inference concerns the ESPN-proxy construction rather than executable market returns. Independence-based win-rate intervals, streak probabilities, Sharpe ratios, and comparisons with live trading are not reported because they would reuse the same correlated rows or compare unlike measurements.

5.1.2 Multiple Testing Considerations

Seven sport models and several thresholds/model families were examined. The artifacts reported here were produced before the trainer default stopped outer-test selection. No formal multiplicity correction or nested outer holdout was used, so neither aggregate nor per-sport rows should be described as confirmatory hypothesis tests.

5.1.3 Edge Stability Over Time

The published artifact does not retain timestamps needed to reproduce the earlier monthly table. The earlier edge-decay and 6-12-month half-life statements are withdrawn. A future stability analysis must use timestamped observations, game-clustered uncertainty, and a pre-specified temporal test.

5.1.4 Comparison to Baselines

The artifact directly compares the model with ESPN WP; it contains no independent market-price series with which to reproduce the earlier random-entry or ESPN-versus-market trading baselines. Those trading-baseline claims are withdrawn. Forecast quality should instead be compared with ESPN using proper scoring rules on a game-grouped untouched holdout.

5.2 Historical Live-Trading Snapshot (March-April 2026)

Metric Value
Total bot-attributed trades 90
Resolved 88
Win rate 62.5%
Net P&L +$67.59
Open positions 2

By Bot/Sport:

Bot Trades Record P&L
Moneyline WP (NBA/MLB/NCAA) 35 22W-12L +$27.26
CS2 (Counter-Strike) 28 13W-14L +$0.42
Tennis (ATP/WTA) 14 10W-4L +$33.23
LoL (League of Legends) 12 9W-3L +$6.18
Soccer (EPL/LIGUE1) 1 1W-0L +$0.50

This table preserves the paper's contemporaneous snapshot; it is not the current public record. It must not be compared numerically with §5.1 because one is realized, size-weighted ledger P&L and the other is a one-contract ESPN-proxy simulation. Current confirmed-position results, including measured execution-identifier coverage rather than a blanket on-chain claim, are published at https://zenhodl.net/results.

Live trading covers additional sports beyond those described in Section 2. CS2 uses an Elo + binomial series model with HLTV live game data. LoL uses an Elo + binomial series model with LoLEsports API data. Tennis uses a hierarchical point-game-set-match probability model with ATP/WTA Elo ratings. Soccer uses a Poisson goal model with Elo-adjusted scoring rates. Full model descriptions for these sports are outside the scope of this paper.[^2]

[^2]: The backtest in Section 5.1 covers only the 7 ESPN-based sports. The live results include esports and tennis models that use different data sources and model architectures.

The small, differently sized and differently composed historical live sample does not validate or invalidate the §5.1 simulation. Any current conclusion must use the confirmed live ledger and an evaluation protocol specified before examining its outcomes.

5.3 Closing-Line Evidence — Retraction (2026-07-21)

Earlier versions reported an 88.6% versus 11.3% winner/loser split after grouping settled trades by final-close CLV. That analysis is retracted. The "close" was the last non-terminal price near settlement, so it incorporated post-entry score changes and eventual outcome: winners mechanically approached 100 and losers 0. Conditioning on its sign and then testing win rate is outcome leakage, not independent evidence.

The control and reproducible audit are published at https://zenhodl.net/clv-evidence. Earlier per-sport significance, failure-mode attribution, and operational conclusions derived from that final-close split are withdrawn here.

The /clv page may remain a descriptive ledger scorecard, but final-close CLV is not used in this paper as proof of model skill. Future execution evidence must use outcome-blind, fixed-horizon executable marks (for example T+60, T+180, and T+300), record source timestamps, and pre-specify exclusions for intervening score events.


6. Limitations and Risks

6.1 Backtest vs Live Performance Gap

The 2,625-row artifact is an ESPN win-probability proxy simulation, not a replay of an exchange order book. Its fixed 3.0c post-proxy cost assumption does not model depth, queue position, fill probability, latency, or adverse selection. It therefore cannot establish executable returns on Polymarket, Kalshi, or any other venue. A market-valid evaluation requires timestamp-aligned native quotes or depth, an explicit fill rule, and outcome-blind markouts.

6.2 Small Live Sample

Section 5.2 is a dated historical snapshot with limited sample size and is preserved for reproducibility, not as the current record. Current confirmed live-position results are maintained at https://zenhodl.net/results; any formal evaluation should freeze its population and protocol before observing outcomes.

6.3 Model Degradation

Market efficiency tends to improve over time as more participants adopt quantitative approaches. The edge we observe may decay as: - More automated trading systems enter prediction markets - Market makers improve their pricing algorithms - Information transmission speed increases

Retraining cadence alone is not a remedy for degradation. A challenger should be triggered by a pre-specified schedule or drift diagnosis, evaluated on an untouched future period, and promoted only after its artifact contract and shadow criteria pass.

6.4 Regime Change

Structural changes in sports (rule modifications, season format changes) or markets (fee structure changes, regulatory actions) can invalidate historical patterns. The system's reliance on ESPN game state data means it is vulnerable to API changes or outages.

6.5 Execution Risk

Production uses price-capped Fill-and-Kill orders, for which partial fills are accepted. Exchange, network, book-depth, stale-quote, and adverse-selection risks remain. The public ledger's measured execution-identifier coverage must not be described as blanket transaction-level or on-chain verification.

6.6 Threats to Validity

Internal validity: The temporal split reduces direct look-ahead, but the historical artifacts selected among model families using the outer test set. Edge thresholds and other choices were also informed by overlapping analysis. The corrected trainer default prevents that selection in future runs; a confirmatory revision still needs nested rolling validation, with every design choice confined to inner folds and an untouched outer test period.

External validity: The historical artifact contains no exchange-native price series, so it is not evidence about Polymarket, Kalshi, PredictIt, or another venue's market microstructure. Generalization across venues, seasons, or leagues is not established.

Construct validity: In §5.1, "edge" is model probability minus an ESPN-proxy ask, not model probability minus an executable market ask. It measures disagreement with that proxy under the simulation's rules; it is not itself a tradable edge estimate.


7. System Architecture

The system consists of: 1. Data Pipeline: Async ESPN polling (5s intervals) + Polymarket WebSocket (real-time) + multi-venue OddsAPI polling 2. Model Layer: when this paper was written, it comprised 7 sport-specific XGBoost/LR ensemble models with isotonic calibration 3. Signal Engine: Edge detection with configurable thresholds, uncertainty gates, and staleness checks 4. Execution Layer: Polymarket CLOB via py-clob-client with FAK (fill-and-kill) orders on Polygon 5. Monitoring: Circuit breakers, feed-quality checks, and reconciliation agents

The API, capture, and unified trading processes run as separately monitored services. Deployment topology and enabled sports change over time; the active configuration and service manifests, not this paper, are the operational authority.


8. Conclusion

This paper documents a forecasting and execution pipeline; it does not establish proven executable alpha. The 2,625-row historical result is an ESPN win-probability proxy simulation. Its game-clustered confidence interval includes zero, and its fixed friction assumptions do not represent observed market fills. The March-April 2026 live table is a historical snapshot; the current confirmed-position record is maintained at https://zenhodl.net/results.

The earlier final-close CLV winner/loser comparison is retracted in §5.3 because its mark incorporated post-entry game information and outcome convergence. The control is published at https://zenhodl.net/clv-evidence.

The next evidentiary standard is clear:

  • nested rolling validation with a genuinely untouched outer test period;
  • exchange-native L2 replay or actual fills with timestamped, pre-specified execution rules;
  • outcome-blind fixed-horizon markouts and explicit treatment of intervening score events; and
  • game-clustered uncertainty, reported by sport as well as in aggregate.

Until those tests are complete, the proxy result is descriptive evidence about model-versus-ESPN disagreement under stated assumptions—not proof of profitable market execution.


References

  • Arrow, K.J. et al. (2008). "The Promise of Prediction Markets." Science, 320(5878).
  • Borghesi, R. (2007). "The Home Team Weather Advantage and Biases in the NFL Betting Market." Journal of Economics and Business, 59(4).
  • Croxson, K. & Reade, J.J. (2014). "Information and Efficiency: Goal Arrival in Soccer Betting." The Economic Journal, 124(575).
  • Glickman, M.E. (1999). "Parameter Estimation in Large Dynamic Paired Comparison Experiments." Applied Statistics, 48(3).
  • Kaunitz, L., Zhong, S. & Kreiner, J. (2017). "Beating the Bookies with Their Own Numbers." arXiv:1710.02824.
  • Page, L. (2012). "It Ain't Over Till It's Over: Yogi Berra Bias on Prediction Markets." Economics Bulletin, 32(2).
  • Sauer, R.D. (1998). "The Economics of Wagering Markets." Journal of Economic Literature, 36(4).
  • Stern, H.S. (1991). "On the Probability of Winning a Football Game." The American Statistician, 45(3).
  • Wolfers, J. & Zitzewitz, E. (2004). "Prediction Markets." Journal of Economic Perspectives, 18(2).
  • Zadrozny, B. & Elkan, C. (2002). "Transforming Classifier Scores into Accurate Multiclass Probability Estimates." KDD '02.


How to Cite This Paper

If you reference this work in your own research, please use one of the following citation formats:

APA:

Evans, C. (2026). Exploiting short-term inefficiencies in sports prediction markets using calibrated win probability models (Version 1.2, corrected July 2026). ZenHodl Research. https://zenhodl.net/research

MLA:

Evans, Coy. "Exploiting Short-Term Inefficiencies in Sports Prediction Markets Using Calibrated Win Probability Models." Version 1.2, corrected July 2026, ZenHodl Research, zenhodl.net/research.

Chicago:

Evans, Coy. 2026. "Exploiting Short-Term Inefficiencies in Sports Prediction Markets Using Calibrated Win Probability Models" (Version 1.2, corrected July 2026). ZenHodl Research. https://zenhodl.net/research.

BibTeX:

@article{evans2026sportspm,
  title={Exploiting Short-Term Inefficiencies in Sports Prediction Markets Using Calibrated Win Probability Models},
  author={Evans, Coy},
  journal={ZenHodl Research},
  year={2026},
  month={July},
  note={Version 1.2, corrected July 2026},
  url={https://zenhodl.net/research}
}

Disclaimer: Past performance does not guarantee future results. Sports prediction market trading involves risk of loss. This paper presents research findings, not investment advice. The 2,625-row historical result is an ESPN win-probability proxy simulation with fixed assumed costs, not a reconstruction of actual market execution. Live results have limited sample sizes and should not be extrapolated.

Companion evidence — correction record

The dated audit at /clv-evidence retracts the earlier final-close CLV winner/loser comparison. Its August 6 amendment identifies the inverse-side table as an algebraic identity, not an independently measured control. The /clv page remains a descriptive ledger scorecard; final-close CLV is not evidence of model skill in this paper.

Future execution evidence must use outcome-blind, fixed-horizon executable marks and pre-specified score-event exclusions.