← Back to blog

Is the Parlay Leaderboard Skill or Luck? We Pre-Registered the Test

By the ZenHodl team — we run the trading bots this blog writes about, and the qualifying live-position record, including losses, is public with its admission rules at /results.

Correction — 2026-08-14, same day, hours after first publication. We adversarially verify our own studies before trusting them, and the verification pass caught a definition error in this one: our registered "stake" was the dealer's cash funding each parlay mint, not the bettor's. In this RFQ market the dealer is the buyer on >99.99% of fills and the recorded fill price is the dealer-side token's — so every ROI denominator was the wrong side's money, and our "favorite vs longshot" style labels were inverted (what we called a 93c favorite ticket was a 7c longshot ticket). All numbers below are corrected (bettor-side stake = ticket price × shares). The headline survives the correction — ranks persist, profits don't — and is slightly stronger: the top quintile's out-of-sample loss deepens when divided by the money bettors actually put up. The registration document carries a dated amendment (v1.1) recording the error, the corrected definitions, and the robustness checks. Two interpretive sentences from the first version ("styles have persistently different cost structures", "correlation strongest among longshot players") did not survive verification and are gone.

Second correction — same day, evening, after an external review. Three further critiques were verified and accepted. (1) What we compute is a fill-to-settlement markout normalized by max-loss capital, not cash ROI — the tape cannot distinguish a fresh mint from an exit of existing inventory, so we now say markout throughout. (2) Settlement censoring: 42% of second-half fills were still unresolved at the run (vs 3% of half 1), and losers resolve earlier than winners by construction — so second-half levels are unreliable and we no longer quote them as returns. (3) The quintile comparison now has the formal test it lacked: the top-minus-bottom difference is +6.7pp with a clustered CI of [−53%, +56%] — this window is uninformative on that question. The September run adds a settlement-lag buffer, a frozen input manifest, and a design hash published before execution.

Every prediction-market leaderboard has the same flaw: it's a list of people who already won. With thousands of wallets trading, the top of any profit ranking will be populated even if nobody has skill — that's what tails of random distributions do. The only honest way to ask "is the leaderboard skill?" is to ask whether the same wallets keep winning out of sample.

We're in an unusual position to answer that for parlays. Since August 6 we've been recording Polymarket's invisible combo market directly from the settlement layer — every fill, every wallet, every outcome, with each parlay's legs decoded from the contract itself. Unlike a leaderboard, the tape includes every loser. There is no survivorship bias to fight, because nobody is missing.

So we pre-registered a persistence test — design locked, thresholds chosen, interpretation rules written down — before computing a single wallet's P&L. (Precisely: locally prespecified — file-order provenance, git-pinned after the fact. The September run's design hash will be published before execution, the same pattern as our on-chain NBA-playoffs pre-commit.) Here's the design, and the corrected first read.

The registered design

First read (corrected): ranks persist — profit shows no sign of following

From 93,742 settled real-party fills, 531 wallets qualified.

The rank correlation is consistent. Spearman ρ = +0.14 between half-1 and half-2 markout, permutation p = 0.002 — with the caveat that the permutation treats wallets as independent, which shared slates violate; the slate-aware null is part of the September design. Across 531 wallets, where you rank in one period predicts where you rank in the next. It also survives the obvious mechanical objection: excluding every fill on parlays a wallet held open across both halves (so no single settlement can touch both sides of the comparison) raises the correlation slightly rather than killing it.

But the top of the leaderboard shows no sign of profit. Half-1 top-quintile wallets' second-half markout was −5.5% of max-loss capital (clustered CI [−15.6%, +8.7%]); the bottom quintile +8.0% with a CI spanning [−32%, +89%]. The formal difference test is +6.7pp, CI [−53%, +56%] — this window is uninformative on the top-vs-bottom question, and second-half levels carry a censoring caveat besides: 42% of second-half fills were still unresolved at the run, and losing parlays resolve earlier than winners by construction. What we can say: the registered detection criterion is not met, and nothing in the data suggests the top of the ranking was making money out of sample. (One more honesty note: qualifying requires activity in both halves — a real copy-trader at the split couldn't know who keeps trading.)

Both things are true at once, and the reconciliation is the actual finding: what persists is how you bet, not whether you beat the market. We measured style directly this time: the rank correlation of a wallet's mean ticket price across halves is +0.74 — favorite-stackers stay favorite-stackers, longshot players stay longshot players. Rank persistence in markout at +0.14 alongside style persistence at +0.74 looks like stable identity, not stable edge: markout persistence is positive in all three style terciles, with differences between terciles inside the noise.

Two more numbers for context:

What we are not claiming — yet

This is the registered exploratory read, and the registration says so in bold: an 8-day window contains few independent game slates, shared legs correlate wallet outcomes, and a positive ρ here can be style clustering rather than ability. No skill claim is being made from this run in either direction.

The confirmatory run is pre-scheduled at ≥30 days of capture (~September 6), with the detection bar written down in the registration verbatim: rank persistence (ρ > 0, permutation p < 0.05) and the top quintile's clustered CI lower bound above the bottom quintile's upper bound. The confirmatory run also adds the controls this run lacked: true settlement-day clustering from the decoded legs, a slate-aware null, and a per-wallet leg-overlap control. We'll publish it either way — including a null. Our track record on publishing unfavorable results is the archive itself.

Why this matters if you trade

If you're copy-trading a parlay leaderboard, the first read says the ranking you're following encodes style, not edge — and the top of the ranking showed no out-of-sample profit signal once you divide by capital actually at risk — though this short window cannot settle the question either way. The way to find out whether anyone in this market is genuinely skilled is a persistence test with the losers included, and that requires the complete tape. We'll keep running it as the data accumulates.

The data

The underlying panel is the same one behind the parlay study: free sample · schema & methods · settlement notes. Wallet-level results are published in aggregate only. Full historical panel, custom cuts, or an anonymous benchmark of your own wallet against the population: [email protected].

ZenHodl records prediction-market microstructure continuously across both major U.S. venues. We pre-register our tests, adversarially verify our own published numbers, and correct loudly on the same day when we find an error — that's the product.

Related reading

Get ZenHodl Weekly

One weekly email with live results, one model insight, and product updates.

Tuesday mornings. No spam.

Want the data behind this post?

Historical sports prediction-market datasets with measured coverage, documented schemas, and disclosed gaps.