Planning a prediction market API integration
Most of the cost of pulling prediction market data is not the HTTP call. It is identifiers, conventions and resolution handling. Decide these before you write the client.
Each venue publishes its own developer documentation and its own terms for programmatic access. Read those first and respect them: rate limits, attribution and permitted use are set by the venue, not by a guide. What follows is the design work that sits on top, and that is the same whichever venues you pull from.
1. Model the identifier properly
Venues use different identifier shapes, and some expose a hierarchy: an event or series that contains several individual contracts. Store the native identifier exactly as returned, alongside the venue, and make the pair your primary key. Never key on the human-readable title. If you want a cross-venue concept of "the same question", that must be a separate mapping table you own, with a confidence level and a reviewer, because it is a judgement rather than a fact.
2. Normalise prices once, at the boundary
Presentation differs between venues: some quote in cents, some in decimals between zero and one. Pick one internal representation, convert at ingestion, and store the original value too so you can prove what the venue actually returned. Record which side of the book a number came from. A midpoint, a last trade and a best bid are three different numbers, and mixing them silently is a source of nonsense in downstream analysis.
3. Plan for resolution, not just for quotes
Live prices are the easy part. The fields that matter later are status, close time, settlement time and the resolved outcome. Poll for status transitions and append a record when a contract resolves, with the timestamp you observed it. Keep the outcome and observation time alongside the forecast. Historical coverage varies, so recording observations as you go makes later evaluation easier to audit. Venues do occasionally correct a resolution, so keep the earlier record and mark which version is current rather than overwriting.
4. Respect pagination and rate limits from day one
- Implement cursor or page handling before you have a backlog, not after a partial sync silently truncates a category.
- Back off on 429 and 5xx responses with jitter, and treat a rate limit as a normal condition rather than an error to alert on.
- Cache aggressively for anything you display repeatedly, and keep the cache age visible in your own interface.
- Log request metadata and the response status for every sync, so a gap in your data has an explanation.
- Redact authorisation headers, API keys and any personal data before anything is written to your logs.
5. Storage shape
| Table | Key fields | Notes |
|---|---|---|
| contract | venue, native_id, title, status, close_at, settlement_source | One row per contract per venue; native_id stored verbatim |
| quote | venue, native_id, observed_at, value, value_kind, source_side | Append only; value_kind records the original convention |
| resolution | venue, native_id, resolved_at, outcome, observed_at, version, is_current | Append-only observations; retain corrections and identify the latest version |
| question_map | question_id, venue, native_id, confidence, reviewed_by | Your own cross-venue judgement, kept separate from venue facts |
6. Decide what you will not do
Two boundaries are worth setting explicitly. First, research systems and execution systems have different reliability requirements; if you are building for research, say so and do not let the pipeline creep toward order placement. Second, an automatic cross-venue match should not be presented to your users as a verified equivalence unless a human has read both sets of rules.
Reference documentation
Field names, authentication and limits are defined by each venue, so check the official documentation before writing against it: docs.kalshi.com and docs.polymarket.com. Riddle is not affiliated with these venues.
Where Riddle fits
Riddle covers the first slice of this work: candidate market discovery across the venues we cover, the contract metadata needed to tell those candidates apart, and the observations Riddle has recorded over time. The v1 API is available under development access, with keys you create yourself in the app. If this is the layer you would rather not build yourself, read the developer overview and we will work through it with you.
Keep reading
Do this work in one place
Riddle brings Kalshi, Polymarket and Limitless onto one screen so the steps above start from something organised.
This guide is general information about how prediction market data is structured. It is not financial, investment or trading advice, and it is not a recommendation about any contract or venue. Always read the venue rules for the contract you are studying.