← Back to blog

Why Your Sports Betting Model Loses Money (Calibration Beats Accuracy)

updated 2026-10-08 calibration machine-learning python tutorial intermediate

By ZenHodl. Dataset documentation, research and model evaluations are linked in the article. A separate filtered ledger of bot-attributed trades, including losses and its admission rules, is public at /results.

Your model predicts 70% win probability. The team wins 58% of the time when you say 70%. That gap is a calibration problem. Its trading effect depends on the prices and costs of the selected positions.

Accuracy measures how often you are right. Calibration measures whether your probabilities mean what they say. For sports betting, calibration helps evaluate probability estimates; profit also depends on the acquisition price, costs and the population you trade. See scikit-learn's calibration guide for the standard techniques.

Why Calibration Beats Accuracy

Two models evaluate a game where the market asks 65 cents:

In this hypothetical example, Model A has a 13c gross expected margin at the 65c price, while Model B has only 1c before fees. Neither wins every trade, and the example does not report realized profits. The danger is that Model B may have a better AUC — it just cannot price contracts correctly. The same trap shows up in feature selection: a more accurate model can make less money than a simpler one.

Measuring Calibration: Brier Score

import numpy as np

def brier_score(y_true, y_prob):
    """Lower is better. 0 = perfect, 0.25 = coin flip."""
    return np.mean((y_prob - y_true) ** 2)

Brier score reflects calibration, discrimination and outcome uncertainty. Compare it with an appropriate baseline on the same held-out population; universal 0.20/0.18 cutoffs do not establish model quality.

Measuring Calibration: ECE

Expected Calibration Error directly measures the gap between predicted probability and observed frequency:

def expected_calibration_error(y_true, y_prob, n_bins=10):
    bin_edges = np.linspace(0, 1, n_bins + 1)
    ece = 0.0
    for i in range(n_bins):
        upper = (y_prob <= bin_edges[i + 1]) if i == n_bins - 1 else (y_prob < bin_edges[i + 1])
        mask = (y_prob >= bin_edges[i]) & upper
        if mask.sum() == 0:
            continue
        bin_conf = y_prob[mask].mean()
        bin_acc = y_true[mask].mean()
        ece += mask.sum() * abs(bin_acc - bin_conf)
    return ece / len(y_true)

ECE is a bin-count-weighted average of absolute frequency discrepancies. It is not a bound that puts every prediction within that many points of reality. Report bin definitions, counts, reliability curves and the evaluation population alongside the scalar.

Visualizing It

A reliability diagram plots predicted probability vs observed frequency. Perfectly calibrated models fall on the diagonal. If your curve bows above, you are under-confident (your 60% predictions win 70%). If below, you are over-confident. Both are fixable.

Predicted probability Observed win rate 01 perfectly calibrated over-confident model Illustrative — not a specific model's data

Fixing Calibration: Isotonic Regression

Isotonic regression learns a monotonic mapping from raw model output to calibrated probabilities:

from sklearn.isotonic import IsotonicRegression

# CRITICAL: fit on validation set, never training data
calibrator = IsotonicRegression(out_of_bounds="clip")
calibrator.fit(raw_probs_val, y_val)

# Apply to new predictions
calibrated_probs = calibrator.predict(raw_probs_test)

If you calibrate on training data, you overfit the calibration curve. Always use a held-out validation set.

Platt Scaling: The Alternative

For smaller datasets, Platt scaling fits a logistic curve:

from sklearn.linear_model import LogisticRegression

platt = LogisticRegression()
platt.fit(raw_probs_val.reshape(-1, 1), y_val)
calibrated = platt.predict_proba(raw_probs_test.reshape(-1, 1))[:, 1]

Choose the calibration method on separate validation data. Isotonic regression is more flexible and can overfit small samples; neither method is universally better for sports betting.

Our Calibration Pipeline

At ZenHodl, calibration is a first-class step:

  1. Train the WP model on training data (seasons 2020-2024)
  2. Generate raw probabilities on a validation set (early 2025)
  3. Fit an isotonic calibrator on the validation set
  4. Apply the calibrator to live predictions
  5. Monitor ECE weekly — retrain if it drifts above 0.04

In backtesting, the calibrator has been measured adding roughly 0.5 cents per trade of edge over the raw model. Correction (2026-09-25): that backtest figure should not be read as a live-profit guarantee, and it compounds into "meaningful profit" only if the edge survives execution costs and holds up in live calibration checks. Measured live, our own bought-side calibration (win rate vs. what we actually paid) is negative in every sport we have enough fills to measure (internal fill study, September 2026). Our CLV retraction and honest aggregate are at /clv-evidence. Calibration makes your stated probabilities honest; it does not by itself guarantee they beat the market.

Common Mistakes

Calibrating on training data. Fitting the calibrator on training predictions can bias evaluation; it does not guarantee a perfect curve or mean the fitted mapping learns nothing.

Ignoring domain shift. A calibrator trained on 2023 NBA data may not transfer to 2026 March Madness. Different sports and seasons require recalibration.

Confusing calibration with discrimination. A model predicting 50% for every game is calibrated only in a population that actually wins 50% of the time; it has no discrimination within that population. You need both.

The Bottom Line

In sports betting, your probabilities are your prices. If you say 78% and the market says 65%, you are buying at 65 cents something you value at 78 cents. If the true chance is 66%, the gross margin at 65c is only 1c and may not cover execution costs. A reliability study makes probability claims easier to inspect; it does not certify profitable execution.


Part of the ZenHodl blog. We write about sports analytics, prediction markets, and building trading bots with Python.

Related reading

Get ZenHodl Weekly

Dataset releases, research notes, and public results, including corrections.

Research and dataset updates from ZenHodl.

Want to build this yourself?

Work through ESPN data collection, Elo, probability modeling, calibration and backtesting in six Jupyter notebooks. Inspect the free first module and course requirements before buying.