pnclENGINEGet API access

CAPACITY NOTES

What to log from an odds poller

Log the transitions, not the polls. Which events deserve a line, which to keep out, and how a small hourly rollup feeds every metric you alert on.

A poller that logs every successful poll fills its log with the one line that tells you nothing. The useful log is a list of transitions: errors, stalls, resumes and restarts, plus a small rollup per hour. Metrics tell you something is wrong; this log tells you why.

Metrics are not logs

The monitoring guide tracks five metrics per board: snapshot age, cursor progress, latency, error rate and request rate. Those numbers answer whether the pipeline is healthy right now. They do not answer what happened at 03:14 when the cursor stalled, and that second question is the one you will actually ask during an incident. The log exists for the second question.

What deserves a line

Log the events that change state, with enough context to act on later:

  • Every 429, with its Retry-After value and the call that triggered it, as the rate limits guide recommends for early warning.
  • Every 401 and 403, immediately and loudly, because those are stop errors, not retries.
  • Every 200 that fails validation, with the envelope shape you actually received.
  • Cursor saves and resumes, including the value you resumed from after a restart.
  • Full board reads, which should be rare enough that each one carries a stated reason.
  • Process starts and stops, so an outage window is provable later.

Each line needs the board, the timestamp and the reason. A line without the board name is a needle in a thirteen-board haystack.

What to keep out

Do not log every successful poll at info level, and never push full response payloads into the event stream. The storage guide already answers the payload question: keep compressed raw responses for a few days if you want to re-derive rows after a parser bug, and query the normalized table. A line per changed selection belongs in the changes table, not in syslog.

The hourly rollup pays for everything

One summary line per board per hour, polls made, events changed, errors by status, feeds both the metrics and the postmortems. From the rollup you can recompute any of the five metrics for any hour you kept, which means the detailed event log can stay short and the dashboard can stay honest.

Log the changes, count the routine, keep the payloads out. The next incident then starts with answers, not with a grep through a million identical lines.