Your model predicts 70% win probability. The team wins 58% of the time when you say 70%. That gap is a calibration problem. Its trading effect depends on the prices and costs of the selected positions.
Accuracy measures how often you are right. Calibration measures whether your probabilities mean what they say. For sports betting, calibration helps evaluate probability estimates; profit also depends on the acquisition price, costs and the population you trade. See scikit-learn's calibration guide for the standard techniques.
Why Calibration Beats Accuracy
Two models evaluate a game where the market asks 65 cents:
- Model A says 78%. Well-calibrated — when it says 78%, the team wins ~78% of the time. Real edge: 13 cents.
- Model B says 78%. Poorly calibrated — when it says 78%, the team wins ~66%. Gross expected margin: 1 cent before costs.
In this hypothetical example, Model A has a 13c gross expected margin at the 65c price, while Model B has only 1c before fees. Neither wins every trade, and the example does not report realized profits. The danger is that Model B may have a better AUC — it just cannot price contracts correctly. The same trap shows up in feature selection: a more accurate model can make less money than a simpler one.
Measuring Calibration: Brier Score
import numpy as np
def brier_score(y_true, y_prob):
"""Lower is better. 0 = perfect, 0.25 = coin flip."""
return np.mean((y_prob - y_true) ** 2)
Brier score reflects calibration, discrimination and outcome uncertainty. Compare it with an appropriate baseline on the same held-out population; universal 0.20/0.18 cutoffs do not establish model quality.
Measuring Calibration: ECE
Expected Calibration Error directly measures the gap between predicted probability and observed frequency:
def expected_calibration_error(y_true, y_prob, n_bins=10):
bin_edges = np.linspace(0, 1, n_bins + 1)
ece = 0.0
for i in range(n_bins):
upper = (y_prob <= bin_edges[i + 1]) if i == n_bins - 1 else (y_prob < bin_edges[i + 1])
mask = (y_prob >= bin_edges[i]) & upper
if mask.sum() == 0:
continue
bin_conf = y_prob[mask].mean()
bin_acc = y_true[mask].mean()
ece += mask.sum() * abs(bin_acc - bin_conf)
return ece / len(y_true)
ECE is a bin-count-weighted average of absolute frequency discrepancies. It is not a bound that puts every prediction within that many points of reality. Report bin definitions, counts, reliability curves and the evaluation population alongside the scalar.
Visualizing It
A reliability diagram plots predicted probability vs observed frequency. Perfectly calibrated models fall on the diagonal. If your curve bows above, you are under-confident (your 60% predictions win 70%). If below, you are over-confident. Both are fixable.
Fixing Calibration: Isotonic Regression
Isotonic regression learns a monotonic mapping from raw model output to calibrated probabilities:
from sklearn.isotonic import IsotonicRegression
# CRITICAL: fit on validation set, never training data
calibrator = IsotonicRegression(out_of_bounds="clip")
calibrator.fit(raw_probs_val, y_val)
# Apply to new predictions
calibrated_probs = calibrator.predict(raw_probs_test)
If you calibrate on training data, you overfit the calibration curve. Always use a held-out validation set.
Platt Scaling: The Alternative
For smaller datasets, Platt scaling fits a logistic curve:
from sklearn.linear_model import LogisticRegression
platt = LogisticRegression()
platt.fit(raw_probs_val.reshape(-1, 1), y_val)
calibrated = platt.predict_proba(raw_probs_test.reshape(-1, 1))[:, 1]
Choose the calibration method on separate validation data. Isotonic regression is more flexible and can overfit small samples; neither method is universally better for sports betting.
Our Calibration Pipeline
At ZenHodl, calibration is a first-class step:
- Train the WP model on training data (seasons 2020-2024)
- Generate raw probabilities on a validation set (early 2025)
- Fit an isotonic calibrator on the validation set
- Apply the calibrator to live predictions
- Monitor ECE weekly — retrain if it drifts above 0.04
In backtesting, the calibrator has been measured adding roughly 0.5 cents per trade of edge over the raw model. Correction (2026-09-25): that backtest figure should not be read as a live-profit guarantee, and it compounds into "meaningful profit" only if the edge survives execution costs and holds up in live calibration checks. Measured live, our own bought-side calibration (win rate vs. what we actually paid) is negative in every sport we have enough fills to measure (internal fill study, September 2026). Our CLV retraction and honest aggregate are at /clv-evidence. Calibration makes your stated probabilities honest; it does not by itself guarantee they beat the market.
Common Mistakes
Calibrating on training data. Fitting the calibrator on training predictions can bias evaluation; it does not guarantee a perfect curve or mean the fitted mapping learns nothing.
Ignoring domain shift. A calibrator trained on 2023 NBA data may not transfer to 2026 March Madness. Different sports and seasons require recalibration.
Confusing calibration with discrimination. A model predicting 50% for every game is calibrated only in a population that actually wins 50% of the time; it has no discrimination within that population. You need both.
The Bottom Line
In sports betting, your probabilities are your prices. If you say 78% and the market says 65%, you are buying at 65 cents something you value at 78 cents. If the true chance is 66%, the gross margin at 65c is only 1c and may not cover execution costs. A reliability study makes probability claims easier to inspect; it does not certify profitable execution.
Part of the ZenHodl blog. We write about sports analytics, prediction markets, and building trading bots with Python.