Observability
Triage a production alert
Turn a pager alert into either a fix already moving or a written all-clear, in the order that finds the cause fastest.
By Toolspoke
Skill procedure
An alert fired. The job is to reach one of two endings within minutes: a cause named and someone working on it, or an all-clear with a sentence saying why the alert was wrong. Anything else is an alert still burning.
Work in this order
- Read the alert itself before anything else. Pull the incident from PagerDuty: what fired, what threshold, on which service, and whether it has fired before. An alert that fires every Tuesday at the same hour is a different problem from one that has never fired.
- Confirm it is real. Hit the service's own health endpoint and look at the last five minutes of request traffic in Grafana or SigNoz. A monitor can be broken while the thing it watches is fine, and saying so early saves everyone else being woken.
- Find what changed. List the deploys of that service in the last two hours — Coolify or Vercel, whichever runs it — and the commits merged since the last good one. Most incidents are the last change, and checking takes thirty seconds.
- Read the errors, not the summaries. Open the newest unresolved Sentry issue for that service and read one real stack trace with its request context. The count tells you the size; only the trace tells you the cause.
- Say what you found, once, in the incident channel. What is broken, who it affects, what you think caused it, and what you are doing next. Four sentences. Post it to Slack on the incident channel, not to a thread nobody is watching.
- Then fix or hand over. If the cause is the last deploy, roll it back — that is its own skill and it is faster than a forward fix. If it is not, keep going, and post again only when the answer changes.
Stop and ask a human
- Before anything that deletes or truncates data, reissues a credential, or scales a production database. These are the destructive tools, and being sure is not the same as being authorised.
- When the fix requires a change to a system this workspace can read but not write.
- When fifteen minutes of looking have produced no cause. Say what you ruled out — that is worth more to the next person than another ten minutes of the same search.
What done looks like
The incident is acknowledged, the channel has a cause or an explicit "no cause yet, here is what is ruled out", and either a rollback is running or a named person has it.