How we propose to grade our own calls.
Verified public results are not available on this page yet. Our forecasting performance is unknown, and we are not publishing an estimate, a sample, or an illustrative figure in its place. Everything below describes the method we propose to use.
These counts are shown as unavailable rather than as a number, because no figure here has been verified for publication. A Brier score computed on a handful of resolutions would say almost nothing, so any future score will carry its resolved count beside it.
Proposed: what would get graded
A forecast is a probability between 0 and 1 published for one specific binary market, for the exact contract as that venue words it. Attention scores, feed rankings and quoted cross-venue differences are not forecasts and would not be graded here.
Proposed: when it would lock
One published forecast per market, timestamped at publication and locked from that point, with the lock time shown beside every scored row.
Proposed: which markets would qualify
Binary markets that resolve on a published rule set and were live on a venue we cover at lock time. Markets that void, cancel or change settlement terms after lock would be listed as excluded, with the reason.
Proposed: how it would be scored
Brier score at resolution: the squared difference between the locked probability and the outcome. Lower is better. 0.00 is a perfect forecast, 0.25 is equivalent to always saying fifty-fifty.
Proposed: the baseline
A score alone means little, so the proposed comparison is the venue bid-ask midpoint observed at the forecast lock time, where available. Beating that baseline would be the test, and failing to beat it would appear in the same table.
Proposed: publication and corrections
The intention is an immutable published record: once a scored row is published it is not deleted or rewritten, and any correction is appended with what changed and when. This process is proposed, not yet implemented, and we are not claiming it is in force today.
What we cover
Kalshi, Polymarket and Limitless. Contracts are surfaced as related when their wording, resolution source and settlement date look close enough to compare, and any such match is subject to your own verification. We do not hold a full orderbook archive, and we do not publish fee-adjusted or size-aware arbitrage.
Corrections
If a published figure or match is wrong, our intended process is to correct it in place and note what changed and when, without deleting a scored forecast. Report an error to hello@getriddle.ai.
When results are published, each row is intended to carry the market, the venue, the locked probability and its lock timestamp, the baseline mid-price, the outcome, and the Brier score for both.
Explore Riddle