← Browse skills

Observability

Investigate an error spike

Work out whether a rise in errors is a new fault, a new user, or a new monitor, and answer in that order before proposing a fix.

By Toolspoke

Skill procedure

A spike is a shape on a graph, and three very different things make that shape. Decide which one before spending an hour on the wrong one.

Separate the three cases

  1. Is the rate up, or is the traffic up? Errors per request, not errors. A launch, a crawler or one enthusiastic integration doubles both and changes nothing about correctness. Read the request count for the same window in Grafana, SigNoz or PostHog before reading the error count.
  2. Is it one fault or many? Group the Sentry issues by culprit for the window. One issue carrying ninety per cent of the volume is a bug. Twenty issues rising together is an infrastructure or dependency problem, and chasing the top one wastes the hour.
  3. Is it new? Check the first-seen timestamp on the leading issue. Something first seen months ago that only now dominates means the rate did not change, the mix did — usually a deploy that removed a louder error, or a monitor whose sampling changed.

Then find the cause

  • Read one full event from the leading issue: the stack trace, the release it was tagged with, and the request parameters. The release tag is the fastest link to a deploy.
  • If it maps to a deploy, read the diff of that deploy over the previous one and look only at the file the trace names.
  • If it maps to no deploy, look outward: a dependency's status page, an expiring certificate, a credential rotated this week, a quota reset date, a table that has just grown past an index's usefulness.

Report it like this

One message: how many users are affected and how, which issue, which release, what caused it, and whether it is getting worse. Put the Sentry link in. Do not attach the graph without the sentence — a graph is not a finding.

Stop and ask a human

  • Before changing sampling rates, muting an issue, or editing an alert threshold. Turning off the measurement is how a spike becomes an outage nobody sees.
  • When the cause is in a third party and the fix is to stop calling them.