First 8 cells from each notebook: ESPN API scraping, NBA summary data, Elo ratings, win probability models, and deployment code.
**Build a Polymarket Prediction Bot from Scratch**
---
In this module, we build a production-grade scraper that pulls **game-state snapshots** from ESPN's public API. By the end, you'll have a dataset of hundreds of thousands of rows — each one a snapshot of a game at a specific moment (score, period, time remaining, ESPN's own win probability) — with the final outcome attached as a label.
This dataset is the foundation for everything else in the course: training win-probability models, calibrating edge thresholds, backtesting strategies, and ultimately running a live bot on Polymarket.
Each row in our final dataset represents one play/moment in a game:
| game_id | sport | home_team | away_team | home_score | away_score | period | seconds_remaining | score_diff | time_fraction | espn_home_wp | home_wins |
|---------|-------|-----------|-----------|------------|------------|--------|-------------------|------------|---------------|--------------|-----------|
| 401584793 | NBA | BOS | MIA | 28 | 22 | 2 | 1380.0 | 6 | 0.479 | 0.712 | 1 |
| 401584793 | NBA | BOS | MIA | 30 | 25 | 2 | 1320.0 | 5 | 0.458 | 0.695 | 1 |
| 401584793 | NBA | BOS | MIA | 30 | 28 | 2 | 1260.0 | 2 | 0.437 | 0.621 | 1 |
**Key columns:**
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install aiohttp pandas numpy matplotlib pyarrow nest_asyncio
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
Install dependencies if needed:
# Uncomment and run if you need to install packages
# !pip install aiohttp pandas pyarrow nest_asyncio
import aiohttp
import asyncio
import json
import time
from datetime import datetime, timedelta, timezone
from zoneinfo import ZoneInfo # stdlib 3.9+; Windows may need: pip install tzdata
from typing import Dict, List, Optional, Tuple
from pathlib import Path
import pandas as pd
import nest_asyncio
# Allow running async code in Jupyter (which already has an event loop)
nest_asyncio.apply()
print("All imports OK")
ESPN has thousands of games across multiple seasons. A normal `requests.get()` loop waits for each response before sending the next — painfully slow.
**Async** sends multiple requests at once. Think of it like ordering 10 pizzas by calling 10 restaurants simultaneously, instead of calling one, waiting for delivery, then calling the next. We'll scrape 60,000+ games in minutes instead of hours.
> **New to async?** Don't worry. Paste any async code cell into Claude or ChatGPT and ask "explain this line by line." The pattern is always the same: `async with session.get(url) as resp`.
---
ESPN exposes a public JSON API that powers their website and mobile app. No authentication is needed — you just hit a URL and get JSON back.
Each sport has a scoreboard endpoint that returns **all games for a given day**:
| Sport | Endpoint |
|-------|----------|
| NBA | `https://site.api.espn.com/apis/site/v2/sports/basketball/nba/scoreboard` |
| NCAAMB | `https://site.api.espn.com/apis/site/v2/sports/basketball/mens-college-basketball/scoreboard?groups=50&limit=300` |
| NHL | `https://site.api.espn.com/apis/site/v2/sports/hockey/nhl/scoreboard` |
| NFL | `https://site.api.espn.com/apis/site/v2/sports/football/nfl/scoreboard` |
| CFB | `https://site.api.espn.com/apis/site/v2/sports/football/college-football/scoreboard` |
| MLB | `https://site.api.espn.com/apis/site/v2/sports/baseball/mlb/scoreboard` |
For a single game's play-by-play and win probability:
```
https://site.api.espn.com/apis/site/v2/sports/basketball/nba/summary?event={game_id}
```
Let's start by fetching one day of NBA games to see the raw structure.
This is 8 of 45 cells. The full module continues with hands-on exercises and working code.
Get This Module Free**Course: Build a Polymarket Prediction Bot from Scratch**
---
In Module 1, we scraped game data from ESPN. Now we need a way to measure **how good each team is** — a single number that captures team strength. That number is the **Elo rating**.
Elo ratings were invented by Arpad Elo in the 1960s for chess, but they work beautifully for any head-to-head competition. The core idea:
The beauty of Elo is that it's **self-correcting**. A team that keeps winning will see their rating rise until the model correctly predicts they'll win — then the updates shrink. It converges on the true strength.
For prediction markets, Elo gives us something concrete: if Team A has Elo 1650 and Team B has Elo 1450, we can compute the **exact probability** that Team A wins. If the market says 55% but our Elo model says 70%, that's a 15-cent edge.
By the end of this module, you'll have:
1. A working Elo rating system from scratch
2. Ratings for every NBA team (and the framework for any sport)
3. Validation that higher Elo actually predicts wins
4. Saved ratings ready to use in the win-probability model (Module 3)
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install pandas numpy matplotlib
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
# ── Quick Start: Sample Data ──────────────────────────────────────────────────
# If you haven't completed the previous module, uncomment and run this cell
# to load sample data so you can follow along without being blocked.
#
# Generates synthetic game results so you can build Elo without running Module 1
# import pandas as pd, numpy as np
# np.random.seed(42)
# teams = [f"Team_{i}" for i in range(30)]
# games = []
# for season in ['2023-24', '2024-25']:
# for _ in range(500):
# h, a = np.random.choice(teams, 2, replace=False)
# home_win = np.random.random() < 0.58 # ~58% home win rate
# games.append({'season': season, 'home_team': h, 'away_team': a, 'home_wins': int(home_win)})
# game_results = pd.DataFrame(games)
# print(f"Loaded {len(game_results)} synthetic games")
---
The entire Elo system comes down to two equations.
Given two teams with ratings $R_A$ and $R_B$, the expected score (win probability) for team A is:
$$E_A = \frac{1}{1 + 10^{(R_B - R_A) / 400}}$$
This is a **logistic function** scaled so that a 400-point Elo advantage gives you a ~91% win probability. Some intuition:
| Elo Difference | Win Probability |
|:-:|:-:|
| 0 | 50.0% |
| +100 | 64.0% |
| +200 | 75.9% |
| +400 | 90.9% |
| -100 | 36.0% |
| -200 | 24.1% |
After the game, we update the rating:
$$R_{\text{new}} = R_{\text{old}} + K \times (S - E)$$
Where:
The K-factor is the single most important tuning parameter:
Think of K as a **learning rate**. Too low and the system is slow to recognize a team got better. Too high and one upset throws everything off.
**Analogy:** Think of Elo like a GPA for sports teams. Beating a top-ranked team is like acing a hard class — your GPA jumps. Losing to a weak team is like failing an easy one — it drops a lot. Over time, every team's "GPA" converges to their true strength level.
The magic number **400** in the formula controls sensitivity. With 400, a team rated 200 points higher wins ~76% of the time. This was tuned for chess and works surprisingly well for sports too.
def update_elo(rating_a: float, rating_b: float, winner: str, k: float = 20.0) -> tuple:
"""
Update Elo ratings after a game.
Parameters
----------
rating_a : float
Current Elo rating of Team A.
rating_b : float
Current Elo rating of Team B.
winner : str
'A' if Team A won, 'B' if Team B won.
k : float
K-factor (learning rate). Default 20.
Returns
-------
tuple
(new_rating_a, new_rating_b)
"""
# Step 1: Expected score for Team A
expected_a = 1.0 / (1.0 + 10.0 ** ((rating_b - rating_a) / 400.0))
expected_b = 1.0 - expected_a
# Step 2: Actual result
actual_a = 1.0 if winner == 'A' else 0.0
actual_b = 1.0 - actual_a
# Step 3: Update
new_a = rating_a + k * (actual_a - expected_a)
new_b = rating_b + k * (actual_b - expected_b)
return new_a, new_b
Team A (1500) beats Team B (1600). Team B is the stronger team on paper, so this is an upset.
This is 8 of 51 cells. The full module continues with hands-on exercises and working code.
Get All 6 Modules — $49**Build a Polymarket Prediction Bot from Scratch**
---
This is the core ML module. By the end, you'll have a trained win probability (WP) model that:
1. Takes the current game state (score, time remaining, team strength)
2. Outputs **P(home team wins)** as a calibrated probability
3. Compares that probability to the Polymarket price to find **edges**
On Polymarket, moneyline contracts trade between $0.01 and $0.99. If a contract for "Lakers win" is trading at $0.63 (63 cents), the market is saying the Lakers have a 63% chance of winning.
But what if your model says 72%? That's a **9-cent edge**. You buy at 63 cents, and over many trades, you collect that edge.
The math:
That's the entire business model. Build a better model than the market, exploit the difference, and let the law of large numbers work for you.
**What is a win probability model?**
Imagine you're watching a basketball game. Home team leads by 12 with 8 minutes left in Q3. What's their chance of winning? You intuitively estimate ~80%.
A WP model does this mathematically. It takes the game state (score difference, time remaining, period, team strength) and outputs a probability between 0% and 100%.
**Why this matters for trading:** Prediction markets sell contracts at a price that reflects the market's estimate (e.g., 72 cents = 72% probability). If your model says 82%, that's a 10-cent edge. Buy at 72c, hold to settlement, win at 82% rate = profit.
**Analogy:** The WP model is like a calculator that converts "game situation" into "fair price." The market gives you the store price. When your calculator says the item is worth more than the store price, you buy.
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install pandas numpy scikit-learn xgboost matplotlib
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
# ── Quick Start: Sample Data ──────────────────────────────────────────────────
# If you haven't completed the previous module, uncomment and run this cell
# to load sample data so you can follow along without being blocked.
#
# Generates synthetic WP training data so you can train models without running Modules 1-2
# import pandas as pd, numpy as np
# np.random.seed(42)
# n = 50000
# df = pd.DataFrame({
# 'game_id': np.repeat(range(n//50), 50),
# 'home_team': np.random.choice(['TeamA','TeamB','TeamC','TeamD','TeamE'], n),
# 'away_team': np.random.choice(['TeamF','TeamG','TeamH','TeamI','TeamJ'], n),
# 'score_diff': np.random.normal(0, 10, n),
# 'seconds_remaining': np.random.uniform(0, 2880, n),
# 'period': np.random.choice([1,2,3,4], n),
# 'home_score': np.random.randint(0, 120, n),
# 'away_score': np.random.randint(0, 120, n),
# 'elo_diff': np.random.normal(0, 100, n),
# 'espn_home_wp': np.random.uniform(0.05, 0.95, n),
# })
# df['time_fraction'] = 1.0 - df['seconds_remaining'] / 2880
# df['home_wins'] = (df['score_diff'] + np.random.normal(0, 5, n) > 0).astype(int)
# print(f"Loaded {len(df)} synthetic training rows across {df['game_id'].nunique()} games")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from pathlib import Path
import warnings
import pickle
warnings.filterwarnings('ignore')
plt.style.use('seaborn-v0_8-whitegrid')
plt.rcParams['figure.figsize'] = (12, 6)
plt.rcParams['font.size'] = 12
# We'll focus on NBA for this module — the same approach works for all sports
SPORT = 'NBA'
print(f"Module 3: Training WP Models for {SPORT}")
print(f"numpy {np.__version__}, pandas {pd.__version__}")
---
In Module 1 we scraped ESPN play-by-play data and saved snapshots as parquet files. Each row is a **moment in a game** — a snapshot of the score, time remaining, and ESPN's own win probability.
If you followed Module 1, your data should be at `espn_wp_data_{SPORT}.parquet`.
> **Want more data?** This module scraped a sample dataset. Run the Module 1 scraper over more dates and seasons to grow it. (Our 25.6-million-row training dataset is no longer sold.)
This is 8 of 45 cells. The full module continues with hands-on exercises and working code.
Get All 6 Modules — $49**Build a Polymarket Prediction Bot from Scratch**
---
This is where most people blow up. They build a model that looks incredible in testing, backtest it, see 75% win rates, go live, and proceed to lose money for three straight weeks.
The problem is almost never the model. It's the backtest.
A bad backtest doesn't just give you wrong numbers -- it gives you *confidence* in wrong numbers. You size up, you run it longer, you double down when the losses start because "the backtest said 73% win rate." By the time you realize the backtest was flawed, you've lost real money.
This module covers:
1. How sports betting backtests differ from standard ML evaluation
2. The hold-to-settlement strategy and its economics
3. Building a rigorous backtester from scratch
4. The five mistakes that make every backtest look amazing (and lose money live)
5. Analyzing results the right way
6. Parameter sensitivity -- finding the real sweet spot vs. overfitting
7. Execution cost modeling -- what your backtest forgets
In ML, you care about accuracy, precision, recall, AUC, Brier score. You split train/test, evaluate, done.
In sports betting, **a model with worse Brier score can make more money**. This is not a paradox -- it's because you don't bet on every game. You only bet when your model disagrees with the market by enough to cover costs. The question isn't "how accurate is the model overall?" It's "how accurate is the model *on the subset of games where it disagrees with the market?*"
A model that's slightly miscalibrated but identifies genuine edges will crush a perfectly calibrated model that agrees with the market on everything.
This means:
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install pandas numpy scikit-learn xgboost matplotlib
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.ticker as mticker
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import SplineTransformer
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
from sklearn.isotonic import IsotonicRegression
import warnings
warnings.filterwarnings('ignore')
plt.rcParams['figure.dpi'] = 120
plt.rcParams['font.size'] = 11
plt.rcParams['axes.grid'] = True
plt.rcParams['grid.alpha'] = 0.3
print("Module 4: Backtesting for Sports Betting")
print("Imports loaded.")
---
On Polymarket, sports contracts resolve to **$0.00 or $1.00**. There is no partial payout. This is fundamentally different from stock trading where you exit at some intermediate price.
**The economics are simple:**
**Expected Value:**
$$\text{EV} = p_{\text{fair}} \times (1 - \text{price}) - (1 - p_{\text{fair}}) \times \text{price} = p_{\text{fair}} - \text{price}$$
If your model says the true probability is **72%** and you buy at **63 cents**:
$$\text{EV} = 0.72 - 0.63 = +0.09 = +9\text{c per share}$$
That's the gross edge. It is additive -- but it is before fees. Polymarket charges a taker fee of `5 * p * (1-p)` cents per share on entry (Section 6), so at 63c that +9c is really about +7.8c. Every EV number in this module is gross until we subtract costs explicitly.
You could try to trade in and out -- buy at 63c, sell at 70c when the game swings. But this introduces:
Hold-to-settlement removes all of these. You only need to answer one question: **"What is the true probability that this team wins?"** If you're right more often than the market, you make money. Period.
# Demonstrate the hold-to-settlement EV math
def calculate_ev(fair_wp, entry_price):
"""Calculate expected value in cents.
fair_wp: model's estimated probability (0-1)
entry_price: what we pay in cents (0-100)
"""
entry_frac = entry_price / 100
ev = fair_wp - entry_frac
profit_if_win = 100 - entry_price
loss_if_lose = -entry_price
ev_check = fair_wp * profit_if_win + (1 - fair_wp) * loss_if_lose
return ev * 100, ev_check # both in cents, should match
# Example scenarios
scenarios = [
(0.72, 63, "Model says 72%, buy at 63c"),
(0.72, 72, "Model says 72%, buy at 72c (no edge)"),
(0.72, 80, "Model says 72%, buy at 80c (negative EV!)"),
(0.55, 45, "Model says 55%, buy at 45c"),
(0.85, 75, "Model says 85%, buy at 75c"),
]
print("Hold-to-Settlement Economics")
print("=" * 70)
print(f"{'Scenario':<42} {'EV (c)':>8} {'Win P/L':>8} {'Lose P/L':>9}")
print("-" * 70)
for fair, entry, desc in scenarios:
ev, _ = calculate_ev(fair, entry)
win_pl = 100 - entry
lose_pl = -entry
marker = " <-- edge" if ev > 0 else (" <-- NO edge" if ev == 0 else " <-- LOSING")
print(f"{desc:<42} {ev:>+7.1f}c {win_pl:>+7.1f}c {lose_pl:>+8.1f}c{marker}")
# Visualize: How edge scales with fair_wp - market_price
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# Left: EV surface
fair_wps = np.arange(0.50, 0.95, 0.01)
entry_prices = np.arange(40, 85, 1)
FW, EP = np.meshgrid(fair_wps, entry_prices)
EV = (FW - EP / 100) * 100 # in cents
ax = axes[0]
c = ax.contourf(FW * 100, EP, EV, levels=np.arange(-30, 35, 5), cmap='RdYlGn')
plt.colorbar(c, ax=ax, label='EV (cents/share)')
ax.contour(FW * 100, EP, EV, levels=[0], colors='black', linewidths=2)
ax.set_xlabel('Model Fair WP (cents)')
ax.set_ylabel('Entry Price (cents)')
ax.set_title('Expected Value per Share')
ax.annotate('Break-even line\n(fair_wp = entry)', xy=(70, 70), fontsize=9,
ha='center', bbox=dict(boxstyle='round', fc='white', alpha=0.8))
# Right: Profit distribution for a realistic edge
ax = axes[1]
np.random.seed(42)
n_trades = 500
fair_wp_sim = 0.72
entry_sim = 63
outcomes = np.random.binomial(1, fair_wp_sim, n_trades)
profits = np.where(outcomes == 1, 100 - entry_sim, -entry_sim)
# Two outcomes, two datasets, two colours. matplotlib wants one colour PER
# DATASET, not per bin — passing four colours for one array raises ValueError.
ax.hist([profits[profits < 0], profits[profits > 0]],
bins=[-65, -60, 35, 40], color=['#d32f2f', '#388e3c'],
edgecolor='white', rwidth=0.6,
label=[f'Loss (-{entry_sim}c)', f'Win (+{100 - entry_sim}c)'])
ax.set_xlabel('Profit per Trade (cents)')
ax.set_ylabel('Count')
ax.set_title(f'Profit Distribution (fair=72c, entry=63c, n={n_trades})')
ax.axvline(x=np.mean(profits), color='blue', linestyle='--', linewidth=2,
label=f'Avg: {np.mean(profits):.1f}c')
ax.legend()
plt.tight_layout()
plt.show()
print(f"\nSimulation: {n_trades} trades at 72% fair / 63c entry")
print(f" Wins: {outcomes.sum()} ({outcomes.mean()*100:.1f}%)")
print(f" Total PnL: {profits.sum():.0f}c ({profits.mean():.1f}c/trade)")
print(f" Theoretical EV: {(fair_wp_sim - entry_sim/100)*100:.1f}c/trade")
---
The backtester simulates exactly what the live bot does:
1. For each in-game snapshot, predict the fair win probability
2. Compare model's estimate to the "market price" (we use ESPN WP as a proxy)
3. If the edge exceeds our threshold AND passes all filters, log a trade
4. The trade resolves to +profit or -loss based on who actually won
**Key design decisions:**
We don't have historical Polymarket orderbook data at second-level granularity. But ESPN publishes a real-time win probability for every game, and Polymarket prices track ESPN WP closely (correlation ~0.95 for NBA/NCAAMB). It's not perfect, but it's the best available proxy for backtesting.
This is 8 of 45 cells. The full module continues with hands-on exercises and working code.
Get All 6 Modules — $49**Build a Polymarket Prediction Bot from Scratch**
---
This is where everything comes together. In the previous modules, we:
1. **Module 1**: Scraped live game data from ESPN's public API
2. **Module 2**: Built Elo ratings to measure team strength
3. **Module 3**: Trained a win probability model on historical game states
4. **Module 4**: Backtested the edge and refined the trading filters
Now we build the **live bot** — the system that runs in real-time during games, detects edges, and places orders on Polymarket.
The bot is a continuous loop with four stages:
```
┌─────────────┐ ┌──────────────┐ ┌───────────────┐ ┌──────────────┐
│ ESPN API │────▶│ WP Model │────▶│ Edge Filter │────▶│ Polymarket │
│ (scores) │ │ (predict) │ │ (min 8c) │ │ (FAK order) │
└─────────────┘ └──────────────┘ └───────────────┘ └──────────────┘
▲ │
│ ┌──────────────┐ │
└────────────────────│ Sleep 5-15s │◀─────────────────────────┘
└──────────────┘
```
This is the key insight: **we never sell**. Every contract on Polymarket resolves to either $0.00 or $1.00 when the game ends. If our model says a contract is worth $0.72 and the market is selling it for $0.62, we buy it and hold until the game finishes.
This eliminates an entire category of risk. The only question is: **is our model better than the market?**
---
**This module can place real orders with real money. It is a method, not a profit
system, and nothing here is a promise that you will make money.**
the system this course is drawn from is **2,722 settled trades, 46.8% win rate,
−$99.73** (2026-03-09 to 2026-09-21). Roughly breakeven to negative. See
`07_live_results_for_notebooklm.md` and <https://zenhodl.net/results>.
$1.00 or $0.00. Only risk money you can afford to lose.
Module 4's own rule, applied to Module 4's own numbers.
spread as a loss. A gap between your fair value and the ask may simply be your model
being under-dispersed.
them before a single live order. The default in this notebook is `mode="shadow"`, and it
should stay that way until you have a settled sample you have actually looked at.
involves substantial risk of loss. **18+ only (21+ where required).** If gambling is a
problem, call 1-800-GAMBLER or visit ncpgambling.org.
everywhere. Check before you trade.
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install aiohttp requests websocket-client py-clob-client
#
# Note the package name: the import in this notebook is `websocket`, which
# comes from `websocket-client`. The similarly named `websockets` package is
# a different library and does not provide it.
# py-clob-client is Polymarket's official client, and the LIVE order path in
# this module depends on it. Every CLOB order must carry an EIP-712 signature
# produced from your wallet's private key — that is not something to hand-roll,
# and a POST without it is rejected no matter how good your HMAC headers are.
# Shadow mode (the default, and where you should spend your first week) needs
# none of it.
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
---
Before we write any code, let's understand the platform we're trading on.
Polymarket is a **prediction market** — a marketplace where you buy and sell contracts on the outcome of real-world events. For sports:
Polymarket uses an on-chain **Central Limit Order Book**, similar to a stock exchange:
| Concept | Description |
|---------|-------------|
| **Bid** | Highest price someone will pay (buy order) |
| **Ask** | Lowest price someone will sell at (sell order) |
| **Spread** | Ask minus Bid — the gap between buyers and sellers |
| **GTC** | Good Till Cancel — order stays on the book until filled or cancelled |
| **FOK** | Fill Or Kill — must fill entirely and immediately, or cancel |
Each side of each market has a **unique token ID** — a long hex string that identifies the contract on-chain:
```
Event: "Lakers vs Celtics"
├── Lakers YES token: 0x1a2b3c... (pays $1 if Lakers win)
└── Celtics YES token: 0x4d5e6f... (pays $1 if Celtics win)
```
In **neg-risk markets** (which cover all sports games), YES + NO always sum to $1.00 at resolution. This means:
Polymarket exposes two APIs:
| API | Purpose | URL |
|-----|---------|-----|
| **Gamma API** | Event/market discovery, metadata | `https://gamma-api.polymarket.com` |
| **CLOB API** | Order placement, orderbook, positions | `https://clob.polymarket.com` |
| **WebSocket** | Real-time price updates | `wss://ws-subscriptions-clob.polymarket.com/ws/market` |
We use Gamma to discover which games are tradeable, CLOB to get prices and place orders, and optionally the WebSocket for faster price updates.
---
Polymarket needs **two different things** from you, and it is worth being clear
about which does what:
1. **API credentials** (key, secret, passphrase) authenticate the *request* —
they prove who is asking. This is the HMAC "L2 auth" you will see below.
2. **Your wallet private key** signs the *order itself* (an EIP-712 signature
over the order struct). Without it there is no order, only a request.
An unsigned POST to `/order` is rejected however well you authenticate it. That
is why live mode in this module uses Polymarket's own `py-clob-client` rather
than a hand-rolled HTTP call — it builds and signs the order for you.
You also need a funded Polymarket wallet (USDC on Polygon) and the funder
address.
Store these in a `.env` file (never commit this to git):
```bash
POLY_PRIVATE_KEY=0xYourWalletPrivateKey # signs each order (live mode only)
POLY_API_KEY=your-api-key
POLY_API_SECRET=your-api-secret
POLY_API_PASSPHRASE=your-passphrase
POLY_FUNDER=0xYourWalletAddress
POLY_SIG_TYPE=0 # 0 = EOA, 1 = email proxy, 2 = browser-wallet proxy
```
**Shadow mode needs none of this.** Spend your first week there.
Before running the live bot, you need a Polymarket account and API credentials.
1. Go to [polymarket.com](https://polymarket.com) and sign up
2. Your wallet is automatically created on the Polygon blockchain
3. Deposit USDC to fund your trading (start with $10-50 for testing)
Install the py-clob-client: `pip install py-clob-client`
Generate your API key:
```python
from py_clob_client.client import ClobClient
client = ClobClient(
host="https://clob.polymarket.com",
chain_id=137, # Polygon
key="YOUR_PRIVATE_KEY",
)
creds = client.derive_api_key()
print(f"API Key: {creds.api_key}")
print(f"API Secret: {creds.api_secret}")
print(f"API Passphrase: {creds.api_passphrase}")
```
```
POLY_PRIVATE_KEY=your_wallet_private_key
POLY_API_KEY=your_api_key
POLY_API_SECRET=your_api_secret
POLY_API_PASSPHRASE=your_passphrase
POLY_SIG_TYPE=0
POLY_FUNDER=your_wallet_address
```
`POLY_PRIVATE_KEY` is the same key you passed to `derive_api_key()` above. It
signs every order, so live mode cannot work without it — and it is exactly the
value you must never paste into anything, commit, or share. If you are not
ready to keep a private key on the machine running the bot, stay in shadow
mode; shadow mode does not read it.
`POLY_SIG_TYPE` must match your account: `0` if you trade directly from the EOA
that owns the key above, `1` for a Polymarket email/magic proxy wallet, `2` for
a browser-wallet proxy. The wrong value produces rejected orders, not an error
you can read.
⚠️ **Never commit your .env file to git!** Add it to `.gitignore`.
import os
import json
import time
import hmac
import hashlib
import base64
import requests
from pathlib import Path
from urllib.parse import urlencode
# ── Load environment variables ──────────────────────────────────────────────
def load_env(env_path=".env"):
"""Load .env file into os.environ."""
p = Path(env_path)
if not p.exists():
print(f"Warning: {env_path} not found. Set POLY_* env vars manually.")
return
with open(p) as f:
for line in f:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
if line.startswith("export "):
line = line[7:]
key, _, val = line.partition("=")
val = val.strip().strip('"').strip("'")
os.environ.setdefault(key.strip(), val)
load_env()
# ── API endpoints ───────────────────────────────────────────────────────────
POLY_BASE = "https://clob.polymarket.com"
GAMMA_BASE = "https://gamma-api.polymarket.com"
POLYGON_CHAIN_ID = 137
# ── L2 request auth (READ endpoints only) ───────────────────────────────────
def clob_auth_headers(method: str, path: str, body: str = "") -> dict:
"""HMAC-signed headers for AUTHENTICATED READ endpoints (balances, your own
orders and trades).
THIS DOES NOT AUTHORIZE A TRADE. L2 HMAC auth proves *who is asking*; it
does not sign the order. An order is a separate EIP-712 signature over the
order struct itself, made with your wallet private key — see
`get_clob_client()` below. Posting an unsigned body to /order fails, always.
Three details that are easy to get wrong and silently fail auth:
1. Header names use UNDERSCORES (POLY_ADDRESS), not hyphens.
2. The secret is URL-SAFE base64 — decode and encode with
urlsafe_b64decode / urlsafe_b64encode. With the standard alphabet,
roughly 60% of real secrets raise binascii.Error and most of the rest
produce different key bytes, i.e. a signature the server rejects.
3. There is no nonce header in the L2 set.
(Verified against py-clob-client 0.34.6: headers/headers.py, signing/hmac.py.)
"""
timestamp = str(int(time.time()))
api_key = os.environ["POLY_API_KEY"]
secret = os.environ["POLY_API_SECRET"]
passphrase = os.environ["POLY_API_PASSPHRASE"]
message = timestamp + method.upper() + path + body
sig = hmac.new(
base64.urlsafe_b64decode(secret),
message.encode("utf-8"),
hashlib.sha256,
).digest()
signature = base64.urlsafe_b64encode(sig).decode("utf-8")
return {
"POLY_ADDRESS": os.environ.get("POLY_FUNDER", ""),
"POLY_SIGNATURE": signature,
"POLY_TIMESTAMP": timestamp,
"POLY_API_KEY": api_key,
"POLY_PASSPHRASE": passphrase,
"Content-Type": "application/json",
}
# ── Signed order client (LIVE mode only) ────────────────────────────────────
_CLOB_CLIENT = None
def get_clob_client():
"""Return an authenticated py-clob-client instance, or raise a clear error.
Raising is deliberate. The previous version of this notebook hand-rolled an
unsigned POST to /order; it could never have filled, and with no .env it
died on an uncaught KeyError *inside the trading loop*, killing the session.
A live path that cannot place an order must say so before it is used, not
fail in the middle of a game.
"""
global _CLOB_CLIENT
if _CLOB_CLIENT is not None:
return _CLOB_CLIENT
try:
from py_clob_client.client import ClobClient
from py_clob_client.clob_types import ApiCreds
except ImportError as e:
raise RuntimeError(
"LIVE mode requires the official client: pip install py-clob-client. "
"Orders must be EIP-712 signed; there is no REST shortcut."
) from e
missing = [k for k in ("POLY_PRIVATE_KEY", "POLY_API_KEY",
"POLY_API_SECRET", "POLY_API_PASSPHRASE")
if not os.environ.get(k)]
if missing:
raise RuntimeError(
f"LIVE mode is missing {missing} in your environment/.env. "
"POLY_PRIVATE_KEY is the wallet key that signs each order; the other "
"three are the CLOB API credentials that authenticate the request."
)
creds = ApiCreds(
api_key=os.environ["POLY_API_KEY"],
api_secret=os.environ["POLY_API_SECRET"],
api_passphrase=os.environ["POLY_API_PASSPHRASE"],
)
# signature_type: 0 = you trade directly from the EOA that owns this key
# (what the production system uses), 1 = Polymarket email/magic proxy,
# 2 = browser-wallet proxy. It must match your account or every order is
# rejected. funder is the address that actually holds the USDC.
_CLOB_CLIENT = ClobClient(
POLY_BASE,
key=os.environ["POLY_PRIVATE_KEY"],
chain_id=POLYGON_CHAIN_ID,
signature_type=int(os.environ.get("POLY_SIG_TYPE", "0")),
funder=os.environ.get("POLY_FUNDER") or None,
creds=creds,
)
return _CLOB_CLIENT
print("API setup complete.")
print(f" Funder: {os.environ.get('POLY_FUNDER', 'NOT SET')[:10]}...")
print(" Shadow mode needs no credentials. Live mode needs POLY_PRIVATE_KEY +")
print(" the three POLY_API_* values, and py-clob-client installed.")
This is 8 of 42 cells. The full module continues with hands-on exercises and working code.
Get All 6 Modules — $49**Build a Polymarket Prediction Bot from Scratch**
---
| Topic | Outcome |
|---|---|
| FastAPI serving | Expose your WP model as a REST API |
| Cloudflare Tunnel | Free HTTPS, no port forwarding, works behind NAT |
| Discord alerts | Real-time edge notifications to your phone |
| Cron + systemd | Automated daily predictions, always-on API |
| Trade logging | JSONL trade log with daily/weekly P&L reports |
| Monitoring | Health checks, stale-data alerts, uptime |
| Scaling | Adding more sports, models, and markets |
| Monetization | Turning your edge into a business |
**Prerequisites:** Modules 1–5 completed. You have a trained WP model, a working Polymarket bot, and historical backtest results.
| Framework | Speed | Auto docs | Async | Type checking |
|-----------|-------|-----------|-------|---------------|
| Flask | Slow | No | No | No |
| Django | Medium | Partial | Partial | No |
| **FastAPI** | **Fast** | **Yes** | **Yes** | **Yes** |
FastAPI is the modern standard for Python APIs. It's async-native (your bot already uses async), generates API docs automatically, and validates request data with type hints. For a trading API that needs to respond in milliseconds, it's the right choice.
**Analogy:** Flask is like a bicycle — simple, gets the job done. FastAPI is like an electric bike — same simplicity, but faster and with more features built in.
# ── Install required packages (run this cell first!) ──────────────────────────
# Uncomment the line below and run if you haven't installed these yet:
# !pip install fastapi uvicorn aiohttp requests
**New to Python? No problem.** Every cell in this notebook is designed to work with AI coding assistants.
If you get stuck on any cell:
1. **Copy the cell** into Claude, ChatGPT, or any AI assistant
2. **Ask:** "Explain this code line by line"
3. **To customize:** "Help me modify this for soccer instead of NBA"
4. **To debug:** Paste the error message and ask "How do I fix this?"
5. **To extend:** "Add a feature that tracks home/away win streaks"
Think of the AI as a patient tutor sitting next to you. The notebooks give you working code — the AI helps you understand and extend it.
> **Pro tip:** If a cell is confusing, ask the AI: "Explain this to me like I've never written Python before." It will break down every line.
---
You have a bot that works on your laptop. It scrapes ESPN, computes fair probabilities, and places trades on Polymarket. That is genuinely impressive — most people never get this far.
But a laptop script is fragile. Your Wi-Fi drops. Your laptop sleeps. You close the terminal by accident. A single uncaught exception kills the process and you miss the best edge of the week.
Production deployment rests on **three pillars**:
Your bot needs to be up when games are on. That means:
Silent failures are the worst kind. Your bot should scream when something goes wrong:
The market evolves. Your models need to evolve with it:
This module builds all three pillars. By the end, you will have a production-grade system that runs itself.
---
Wrapping your WP model in an API has two benefits:
1. **Decoupling** — The model runs as a service. Your bot, your dashboard, your mobile app, and your paying subscribers all hit the same endpoint.
2. **Monetization** — An API with an auth key is a product you can sell.
# Install dependencies (run once)
# !pip install fastapi uvicorn
# ---- api/app.py ----
# Save this as a standalone file. FastAPI apps run via uvicorn, not inside notebooks.
#
# THE MODEL CONTRACT — this is where model servers break.
# Module 3 pickles a BUNDLE, not a bare estimator:
# {"model", "model_type", "sport", "feature_names", "elo_ratings",
# "game_seconds", "train_seasons", "test_season", "metrics"}
# The server must build its feature row from bundle["feature_names"] — same
# names, same order, same arity. Hardcode a row instead and every request dies
# with "X has 6 features, but SplineTransformer is expecting 5" while
# /v1/health still cheerfully reports the model as loaded. The startup
# smoke-predict below turns that whole class of bug into a boot-time failure
# instead of a per-request 500 that nobody is watching.
API_CODE = '''
import os
import glob
import time
import pickle
import numpy as np
from fastapi import FastAPI, Query, HTTPException
from fastapi.middleware.cors import CORSMiddleware
# ---------------------------------------------------------------------------
# Constants
# ---------------------------------------------------------------------------
# Fallback only. The bundle's own "game_seconds" wins when present, so a sport
# you train later does not need an entry here.
TOTAL_SECONDS = {
"NBA": 48 * 60, # 48 minutes
"NCAAMB": 40 * 60, # 40 minutes
"NHL": 60 * 60, # 60 minutes
"NFL": 60 * 60, # 60 minutes
"CFB": 60 * 60, # 60 minutes
}
# ---------------------------------------------------------------------------
# App + CORS
# ---------------------------------------------------------------------------
app = FastAPI(
title="Fair Probability API",
description="Real-time win-probability estimates for live sports games.",
version="1.1.0",
)
app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_methods=["GET"],
allow_headers=["*"],
)
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def unwrap_elo(obj):
"""Module 2 writes a WRAPPER file, not a bare ratings dict:
{"sport", "generated_at", "params", "n_teams", "n_games_processed", "ratings"}
The team ratings live under "ratings". Read the top level by mistake and
every team lookup misses, silently falls back to 1500, and elo_diff becomes
a constant 0 — an Elo-blind server that still reports a healthy-looking
team count (6, the number of wrapper keys). Unwrap once, here.
"""
if isinstance(obj, dict) and isinstance(obj.get("ratings"), dict):
return obj["ratings"]
return obj if isinstance(obj, dict) else {}
def build_feature_row(feature_names, state):
"""Build X in the BUNDLE'S OWN feature order and arity.
A missing feature is a 400 naming the feature — never a silent zero, which
is how a model ends up served with a dead input column.
"""
row = []
for name in feature_names:
if name not in state:
raise HTTPException(
status_code=400,
detail=(
"Model expects feature " + repr(name) + ", which this endpoint "
"does not compute. Computed: " + str(sorted(state)) + ". "
"Add it to the state dict or retrain without it."
),
)
row.append(float(state[name]))
return np.array([row], dtype=float)
# ---------------------------------------------------------------------------
# Load models at startup — and SMOKE-PREDICT every one of them
# ---------------------------------------------------------------------------
models = {}
broken = {}
MODEL_DIR = os.environ.get("MODEL_DIR", "models")
for path in sorted(glob.glob(os.path.join(MODEL_DIR, "wp_model_*.pkl"))):
sport = os.path.basename(path)[len("wp_model_"):-len(".pkl")].upper()
try:
with open(path, "rb") as f:
bundle = pickle.load(f)
estimator = bundle["model"]
feature_names = list(bundle.get("feature_names") or [])
if not feature_names:
raise ValueError(
"bundle has no feature_names — retrain with Module 3 cell 38"
)
# Arity assert: the bundle's declared contract vs what the estimator was
# actually fit on. These disagreeing is the 500-on-every-request bug.
n_expected = getattr(estimator, "n_features_in_", None)
if n_expected is not None and int(n_expected) != len(feature_names):
raise ValueError(
"feature_names declares " + str(len(feature_names)) +
" features but the estimator was fit on " + str(int(n_expected))
)
# Zero-row smoke predict. Cheap, and it is the only thing that proves
# "loaded" means "can actually answer a request".
probe = np.zeros((1, len(feature_names)), dtype=float)
if hasattr(estimator, "predict_proba"):
estimator.predict_proba(probe)
else:
estimator.predict(probe)
except Exception as e:
broken[sport] = type(e).__name__ + ": " + str(e)
print("BROKEN " + sport + ": " + broken[sport] + " — will not be served")
continue
models[sport] = bundle
n_teams = len(unwrap_elo(bundle.get("elo_ratings", {})))
print("Loaded " + sport + ": " + str(len(feature_names)) +
" features, " + str(n_teams) + " teams in Elo table")
if not models:
# Fail loudly at boot. A server with zero usable models that still answers
# 200 on /v1/health is the silent-death shape: nothing pages, nothing works.
raise RuntimeError(
"No usable models in " + repr(MODEL_DIR) + " (broken: " + str(broken) + ")"
)
BOOT_TIME = time.time()
# ---------------------------------------------------------------------------
# Routes
# ---------------------------------------------------------------------------
@app.get("/v1/health")
def health():
"""Health check for uptime monitors.
"Loaded" here means loads AND predicts. A model that unpickles but cannot
predict is reported under models_broken and the service reads degraded.
"""
return {
"status": "degraded" if broken else "ok",
"uptime_s": round(time.time() - BOOT_TIME, 1),
"models_loaded": sorted(models),
"models_broken": broken,
}
@app.get("/v1/predict")
def predict(
sport: str = Query(..., description="Sport code: NBA, NCAAMB, NHL, NFL, CFB"),
score_diff: int = Query(..., description="Home score minus away score"),
seconds_remaining: float = Query(..., description="Seconds left in regulation"),
period: int = Query(..., description="Current period (1-indexed)"),
home_team: str = Query(..., description="Home team key, as spelled in the Elo table"),
away_team: str = Query(..., description="Away team key, as spelled in the Elo table"),
is_home_batting: int = Query(0, description="MLB only; ignored unless the model asks for it"),
):
"""Return fair win probability for a live game state."""
sport_key = sport.upper()
if sport_key in broken:
raise HTTPException(
status_code=503,
detail="Model for " + sport_key + " loaded but cannot predict: " + broken[sport_key],
)
bundle = models.get(sport_key)
if not bundle:
raise HTTPException(status_code=404, detail="No model loaded for " + sport_key)
total_sec = float(bundle.get("game_seconds") or TOTAL_SECONDS.get(sport_key, 2880))
# Range guard, same contract as the bot in Module 5. time_fraction below is
# clamped to [0, 1], but seconds_remaining goes into the model RAW, so a
# clamp there would hide the problem instead of reporting it: an impossible
# clock would still be priced, just quietly. A state the training column
# could never take is refused, not extrapolated. Overtime is the usual cause
# (Module 1 drops OT rows, so the model never saw one).
if not (0.0 <= seconds_remaining <= total_sec):
raise HTTPException(
status_code=400,
detail=(
"seconds_remaining=" + str(seconds_remaining) + " is outside "
"[0, " + str(total_sec) + "] for " + sport_key + ". Training never "
"saw that state, so this endpoint will not price it. If the game is "
"in overtime, do not send it here."
),
)
if period < 1:
raise HTTPException(status_code=400, detail="period must be >= 1")
ratings = unwrap_elo(bundle.get("elo_ratings", {}))
home_elo = float(ratings.get(home_team, 1500.0))
away_elo = float(ratings.get(away_team, 1500.0))
# Every feature this endpoint knows how to compute. The BUNDLE decides which
# of them are used, and in what order.
state = {
"score_diff": score_diff,
"seconds_remaining": seconds_remaining,
"period": period,
"time_fraction": min(max(1.0 - seconds_remaining / total_sec, 0.0), 1.0),
"elo_diff": home_elo - away_elo,
"is_home_batting": is_home_batting,
}
feature_names = list(bundle["feature_names"])
X = build_feature_row(feature_names, state)
fair_wp = float(bundle["model"].predict_proba(X)[0, 1])
return {
"sport": sport_key,
"home_team": home_team,
"away_team": away_team,
"home_fair_wp": round(fair_wp, 4),
"away_fair_wp": round(1 - fair_wp, 4),
"home_elo": home_elo,
"away_elo": away_elo,
# False means one or both names were missing from the Elo table and the
# model was handed a default 1500. Surface it; do not hide it.
"elo_matched": (home_team in ratings) and (away_team in ratings),
"features_used": feature_names,
"score_diff": score_diff,
"seconds_remaining": seconds_remaining,
"period": period,
}
@app.get("/v1/models")
def list_models():
"""List servable models, their feature contract, and anything broken."""
info = {}
for sport, bundle in models.items():
info[sport] = {
"teams": len(unwrap_elo(bundle.get("elo_ratings", {}))),
"features": list(bundle.get("feature_names") or []),
"model_type": bundle.get("model_type") or type(bundle["model"]).__name__,
}
for sport, why in broken.items():
info[sport] = {"error": why}
return info
'''
print(API_CODE)
This is 8 of 40 cells. The full module continues with hands-on exercises and working code.
Get All 6 Modules — $49