RATE LIMITS
Pinnacle API rate limits are a budget, not a wall.
Every Pinnacle odds API plan carries a per-second ceiling. Cross it and the REST API answers 429 instead of data. Handled well, a 429 is a scheduling signal: honor Retry-After, back off with jitter, and keep enough headroom that you rarely see one.
What a Pinnacle API limit response looks like
Exceed the per-second allowance and the API returns HTTP 429 Too Many Requests, normally with a Retry-After header that says how many seconds to wait before the next attempt:
HTTP/1.1 429 Too Many Requests
Retry-After: 2
Content-Type: application/json
Treat 429 as a signal, not an exception. Log every occurrence with its timestamp and the call that triggered it: a rising 429 count is an early warning that steady-state load is drifting toward the ceiling. Retry-After is authoritative. When it is present, it overrides whatever schedule your client had in mind, and hammering through a 429 without waiting is how a transient limit becomes a longer throttle.
Backoff that works
The standard pattern is exponential backoff with jitter: wait a base delay, double it after each rejection, cap it, and add randomness so parallel workers do not retry in lockstep.
delay = 0.5 # seconds
max_delay = 30
loop:
res = get(url)
if res.status != 429:
return res
wait = res.headers["Retry-After"] ?? delay
sleep(wait + random(0, wait / 2))
delay = min(delay * 2, max_delay)
Three rules make it work in production. Always honor Retry-After when the header exists, because the server knows its own window better than your estimate. Cap total attempts and fail the cycle instead of blocking the next one; the odds will be fresher in five seconds anyway. Share backoff state across workers, because ten processes each backing off politely still add up to ten times the requests.
Designing headroom
Backoff is triage; headroom is prevention. Keep steady-state usage under roughly 70 percent of the per-second quota. On a 10 requests per second tier that means pacing regular polling at 7 per second or less, and reserving the rest for retries, admin tooling and a second environment.
Match days are the test. Board requests stay flat when the slate grows, but per-event detail calls, retries and reconnect bursts do not, as capacity planning shows, and the first 429 responses tend to arrive during the biggest slates, when staleness costs the most. If your peak estimate sits at 80 or 90 percent of the ceiling, the answer is a larger tier or a smaller scope, not tighter retries. The tiers compared on this site differ mainly in per-second rate, which is exactly the dimension headroom lives in.
When polling stops scaling
Polling cost grows linearly with freshness: halve the interval and you double the requests, across every board you poll, all day. Past a point the curve stops making sense. If your product needs sub-2-second awareness across a full slate, REST polling is the wrong tool for the hot path.
The alternative is push instead of pull. SSE drop alerts stream processed odds changes as they happen, while REST stays on for bootstrap and periodic reconciliation. The poller asks for a snapshot; the stream tells you what just moved, and per-second limits mostly stop being your problem.
The polling discipline in efficient REST integration still applies to the REST side of that split, and the plan comparison shows which tiers bundle the stream.
A loop that already handles 429 the way this page describes is on the Node.js quickstart on pnclFEED, where Retry-After sets the wait.