← Back to blog

ZenHodl vs Kaggle: Free & Paid Polymarket Data

By the ZenHodl team — we run the trading bots this blog writes about, and the qualifying live-position record, including losses, is public with its admission rules at /results.

This page is for researchers and analysts comparing two very different ways to get Polymarket historical data: free community datasets on Kaggle versus paid, curated archives from ZenHodl. If you need a no-cost dataset for exploration and only need Polymarket, Kaggle may be the right starting point. If you need Kalshi coverage, cross-venue quote alignment, citable provenance, or a curated archive you can own and backtest offline, ZenHodl is the closer match.

At a Glance

Aspect ZenHodl Kaggle Polymarket Datasets
Venues Polymarket + Kalshi, across 15 one-time archive SKUs plus a live API [1] Community Polymarket datasets are common; small, general (non-sports) Kalshi datasets also exist, but nothing at ZenHodl's Kalshi sports order-book scale was found [2]
Time depth Flagship archive: fixed Dec 2025–Jan 2026 window (~136M rows); Kalshi Microstructure Tape: Jun–Sep 2026 [3] One community tick-level dataset covers a 21-day window (Mar 6–26, 2026), now on its 4th revision (updated Aug 2026) with corrected features [2]
Pricing One-time archives $9–$199; per-sport Kalshi slices $39; live API $0–$499/mo; free samples [1][4] Free [2]
Resolution Varies by SKU — the flagship archive's Kalshi side is top-of-book only; genuine L2 depth (up to 20 levels/side) plus trade prints is the separate Kalshi Microstructure Tape and its per-sport slices [3] Tick-level, event-driven orderbook updates per the dataset's own description, not a forward-filled continuous panel [2]
Order-book depth Yes on the Kalshi-specific products — up to 20 levels/side, not sold as a full book; top-150-by-volume-per-cycle sample, not a census [3] Yes, L2 snapshots per the dataset's own description [2]
Score-sync Partial: only Kalshi rows in the flagship archive carry captured score/period/clock at capture time; Polymarket rows do not. The separate MLB Cross-Venue Matched Book aligns both venues by game and capture time [3][5] Not shown on the datasets examined [2]
Free tier Free $0/mo API plan; a free MLB matched-book sample (CC-BY-NC-4.0) on Hugging Face/Zenodo; $9 tryout archives; free per-SKU sample rows [4][6] Free, standard for Kaggle [2]
Format Parquet + CSV, ZIP archives [1] CSV/Parquet, varies by uploader [2]
Kalshi support Dedicated Kalshi depth + trade-print archives, per-sport slices, and a settled-outcomes layer; depth coverage is a top-150-by-volume subset each cycle [3] A small, general (non-sports) Kalshi market dataset exists; no Kalshi sports order-book depth archive at ZenHodl's scale was found as of this check [2]
Citation/DOI Zenodo DOI for the free MLB matched-book sample (10.5281/zenodo.20816908), confirmed resolving [6] No DOI system; each dataset sets its own license (CC-BY-NC-4.0, CC BY 4.0, and Apache 2.0 all seen across datasets checked) [2]
Quality control Curated, with explicit per-product coverage/gap disclosures [3][5] Community-maintained; quality and documentation vary by uploader
Ownership/licensing One-time purchase; download and backtest offline. Standard license covers single-user research; redistribution needs separate commercial terms [1] Free download subject to each uploader's chosen license and Kaggle's platform terms

When to Pick Kaggle

When to Pick ZenHodl

The Key Differences

1. Free community data vs. curated paid archives

Kaggle is genuinely free and community-maintained, with no official Polymarket dataset. Coverage, frequency, and documentation vary by uploader; one tick-level community dataset covers a three-week window and is event-driven rather than a guaranteed-continuous panel [2]. ZenHodl charges $9–$199 across its one-time archives and discloses named gaps, sampling limits, and a top-150-by-volume Kalshi selection rule rather than presenting any archive as a complete order-book census [1][3]. The right choice depends on whether free access or disclosed, curated structure matters more for your project.

2. The Kalshi gap

Kaggle's well-known Polymarket datasets do not cover Kalshi, and a search of Kaggle's own dataset listings on 2026-09-25 turned up only small, general (non-sports) Kalshi datasets — nothing resembling sport-specific order-book depth. ZenHodl sells dedicated Kalshi depth and trade-print archives, sport-specific slices, and a settled-outcomes layer alongside its Polymarket archives [3]. If your backtest requires Kalshi sports order-book data specifically, ZenHodl is the stronger fit on the evidence gathered here; if you only need Polymarket, Kaggle's free datasets may be enough.

3. Game-state context and ownership model

In ZenHodl's flagship archive, only the Kalshi table carries captured game state at capture time — Polymarket rows in that archive do not, and event-time work needs a separate verified join [3]. A separate product, the MLB Cross-Venue Matched Book, aligns Polymarket and Kalshi quotes by dated game key and capture time [5]. Neither is shown in the Kaggle community datasets examined [2]. ZenHodl's archives are one-time purchases you own and can backtest offline without an API key or subscription [1]; Kaggle offers community uploads whose license and long-term availability depend on the individual uploader.

Free Sample & Next Steps

ZenHodl also provides a free sample: a single completed MLB game's matched-book data (Polymarket and Kalshi prices aligned, settled outcome labeled), licensed CC-BY-NC-4.0 and available on Hugging Face and Zenodo, linked from ZenHodl's samples page [6]. If you want to compare Kaggle's free community files against a ZenHodl free sample before paying, start here: zenhodl.net/polymarket-historical-data. For the full archive catalog and current pricing, see zenhodl.net/products.

For Research & Citation

The free MLB matched-book sample is DOI-minted on Zenodo for academic citation (10.5281/zenodo.20816908), confirmed resolving as of this check, and linked from ZenHodl's samples page [6]. The paid archives are delivered as downloadable Parquet/CSV files for offline use; no API key is required. Kaggle datasets are community-maintained; cite them according to each uploader's own license (CC-BY-NC-4.0, CC BY 4.0, and Apache 2.0 were all observed across datasets checked for this page) [2].

Sources (checked 2026-09-25)

  1. zenhodl.net/products and zenhodl.net/pricing — ZenHodl archive catalog, SKUs, and API tiers
  2. Kaggle dataset listings, re-fetched 2026-09-25 via Kaggle's public dataset API (search: "polymarket" and "kalshi"); the tick-level Polymarket orderbook dataset referenced is at v4 (updated 2026-08-04), license CC BY-NC 4.0, ~42.6GB, confirmed unchanged from the prior check
  3. zenhodl.net/kalshi-historical-data and zenhodl.net/polymarket-historical-data — archive resolution, depth, coverage disclosures
  4. zenhodl.net/pricing — API tier names and prices
  5. zenhodl.net/products/mlb_matched_archive — MLB Cross-Venue Matched Book product page
  6. doi.org/10.5281/zenodo.20816908 — Zenodo record for the free MLB matched-book sample (re-confirmed resolving 2026-09-25); also on Hugging Face (re-confirmed live 2026-09-25)

Related reading

Get ZenHodl Weekly

One weekly email with live results, one model insight, and product updates.

Tuesday mornings. No spam.

Want the data behind this post?

Historical sports prediction-market datasets with measured coverage, documented schemas, and disclosed gaps.

Join the community

Discuss strategies, share results, get help.

Join Discord