Curso práctico: datos deportivos, modelos y backtesting
Curso de bots para Polymarket:
construye un bot en Python.
Aprende a recopilar datos de ESPN, construir ratings Elo, estimar probabilidades y evaluar calibración y backtests. Los benchmarks públicos documentan evaluaciones concretas; no garantizan los resultados del modelo que construyas.
6 notebooks de Jupyter. La recopilación cubre NBA, NCAAMB, NHL, NFL y CFB; el notebook de modelos también incluye MLB. Revisa el módulo gratuito y sus requisitos antes de comprar.
Same method evaluated in public benchmarks; our filtered public results ledger includes losses. ver los benchmarks públicos
Pago único de $49 por 6 notebooks. Garantía condicional de 30 días; consulta las condiciones del curso.
The course teaches the method. Want to see it running? Our Edge Finder displays model estimates and available venue prices. Public benchmark artifacts cover defined evaluation windows; the filtered public ledger separately reports recorded outcomes.
Abrir Edge Finder →Gratis para explorar estimaciones del modelo y precios disponibles. La cobertura y la antigüedad de las cotizaciones varían por deporte y fuente. Sin tarjeta.
The model you build in this course is the same kind we run live. Public benchmark artifacts document selected evaluations, and our filtered public results ledger includes wins and losses. The results page separately reports measured execution-identifier and direct-transaction-hash coverage.
Works with AI assistants
An AI coding assistant can help explain a cell or error. Include the relevant definitions and test its suggestions. The notebooks still require Python setup, ordered execution and evaluation; adding a new sport requires additional work.
How to evaluate your model
The course teaches calibration, probability error and ranking on held-out data. Metric values depend on the sport, prediction time, outcome prevalence and test sample; the course does not promise a universal target.
| Metric | What it measures | How to compare |
|---|---|---|
| ECE | Calibration gap | State bins and sample size |
| Brier | Probability error | Compare on the same test rows |
| AUC | Ranks winners | Measure ranking on held-out rows |
| vs sharp books | Pinnacle / Poly | Use a dated paired benchmark |
The linked NBA 2026 playoff benchmarks compare paired pregame forecasts under their published timing and inclusion rules. They do not establish universal market parity, current trading profit or the performance of a student model. Full statistical validation · Benchmark evidence
Public benchmark artifacts document selected model evaluations, while the filtered public results ledger shows resolved wins and losses under its published inclusion rules.
- ✓ Public benchmark artifacts and methodology
- ✓ Filtered public results include losses, not just wins
- ✓ Measured execution-ID and direct-tx-hash coverage disclosed
Start with evidence, not promises.
What you'll build — module by module
Six Jupyter notebooks cover the pipeline from data collection to deployment. The display below is an illustrative excerpt and example output, not a recorded notebook run or a guaranteed training result.
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.calibration import CalibratedClassifierCV
# Train on your scraped games (NBA, NCAAMB, NHL, NFL, CFB)
model = GradientBoostingClassifier(n_estimators=200)
model.fit(X_train, y_train)
# Isotonic calibration for reliable probabilities
calibrated = CalibratedClassifierCV(model, method="isotonic")
calibrated.fit(X_cal, y_cal)
print(f"Brier: {brier_score(y_test, calibrated.predict_proba(X_test)[:, 1]):.4f}")
Brier: 0.1395 (illustrative value)
ECE = how far our stated probabilities sit from reality (lower is better); AUC = how well the model ranks winners; Brier = overall probability error.
Scraping ESPN
- Build an async data scraper for ESPN play-by-play
- Handle rate limits, retries, and API pagination
- Collect available historical games across NBA, NCAAMB, NHL, NFL and CFB
- Store as efficient Parquet files
Historical data collection in one notebook
Try free ↓ · AI prompt: "Explain how this async scraper works"
Elo Ratings
- Implement Elo from scratch (no libraries)
- Home advantage, K-factor tuning, season resets
- Team-keyed ratings for the included sports
- Validate against known rankings
The simplest feature that matters most
AI prompt: "Help me add season-decay to my Elo system"
WP Models
- Train LR+Spline and XGBoost+Isotonic
- Isotonic calibration for probability accuracy
- Why a simpler, well-calibrated model often ranks best
- Sport-specific feature engineering
Measure probability error on held-out games
AI prompt: "Explain isotonic calibration like I'm a beginner"
Backtesting
- Time-split validation (train on N-1, test on N)
- Adverse selection and underdog traps
- Execution cost modeling (fees, slippage)
- Deduplication and subsampling
Check leakage, duplicates and execution assumptions
AI prompt: "Help me add a new sport to this backtest"
Live Bot
- Polymarket CLOB API integration
- ESPN adaptive polling (5s/15s)
- Comparing fair value to live market prices
- Order execution with shadow mode (paper-first)
Run it in shadow mode before risking a dollar
AI prompt: "Help me add Discord alerts to this bot"
Deployment
- FastAPI server with HTTPS
- Discord alert webhooks
- Cron scheduling and monitoring
- Cloudflare Tunnel for secure access
From laptop to 24/7 production
AI prompt: "Help me deploy this to a $7/mo VPS"
Want to see the actual code before buying?
Preview the first 8 cells of every module — real teaching, real code, not marketing copy.
Try Module 1 free
Scraping ESPN: build an async historical-data pipeline with retries and Parquet output. Enter your email to download the full Module 1 notebook.
Check your email for the download link too.
No spam. Unsubscribe anytime.
Before and after
Before
- "I don't know Python but want to build a bot"
- Watching tutorials that stop at theory
- Buying picks from Discord that lose money
- No systematic edge, just gut feel
- No way to measure if a strategy works
After
- Your own probability-model pipeline to train and evaluate
- AI-assisted coding — paste any cell and ask for help
- A dataset and evaluation workflow you can inspect
- The skill to measure a model honestly (ECE, Brier, AUC)
- A backtest framework that tells you the truth before you risk a dollar
Built by a trader, not a marketer
These notebooks teach the data, modeling, backtesting and deployment workflow behind ZenHodl. Production models and risk rules evolve separately. The public benchmarks document specific evaluation windows, and the filtered public results ledger reports recorded outcomes under its inclusion rules.
I built this because I couldn't find a prediction market course that showed real code and measured itself honestly. I wanted a course that continued beyond a confusion matrix. This one starts at the data pipeline and ends with a calibrated model you can deploy and audit yourself.
See live results at /results. Read the methodology at /methodology. Read our research paper.
Frequently asked questions
What do I need to get started? +
What skill level is this for? +
What you DO need: A computer (Mac, Windows, or Linux), an internet connection, and the willingness to follow instructions and experiment. You will still need to follow setup steps and check the results.
Is there a refund policy? +
The public benchmarks document specific model versions and evaluation windows; they do not guarantee the performance of a student bot. You can try Module 1 completely free before buying to make sure you like the teaching style.
Note: Refund requires demonstrated completion of all modules. Dataset purchases are non-refundable. Backtests are historical and optimistic — live results run lower and you can lose money. The course teaches the method; it is not a guarantee of profit.
How do I use AI with this course? +
- Copy the cell into Claude, ChatGPT, or any AI assistant
- Ask: "Explain this code line by line" or "What does this function do?"
- To customize: "Help me modify this to track soccer instead of NBA"
- To debug: paste the error message and ask "How do I fix this?"
Use AI explanations as suggestions, then test them against the notebook and your output. Extending to a new sport requires its own data, features and evaluation.
How long does it take to complete? +
Will this work for my sport? +
What's the difference between the course and the Edge Finder? +
What can go wrong
Model drift. Markets adapt. A calibrated fair-value reference can become stale and may need retraining as market efficiency improves.
Execution costs. Use the cost assumptions stated by each notebook or benchmark. Actual fees, slippage and available liquidity depend on the venue and order.
Small samples. Backtests are historical and optimistic. Our filtered public results ledger, including losses, is on /results — a model being well-calibrated does not mean it beats the market or turns a profit.
Regime change. Rule changes, new market participants, or structural shifts can invalidate historical patterns.
Past performance does not guarantee future results. This is a tool for informed decision-making, not a guaranteed profit machine.
Build a model you can trust
6 notebooks. Working code. The same method evaluated in our public benchmark artifacts.