Observability
Confirm a reported outage
Check the whole path between a customer and the thing they say is down, so the reply is a fact rather than a reassurance.
By Toolspoke
Skill procedure
Somebody says it is down. Before replying anything, find out. The reply that costs the most trust is "it looks fine on our side" sent without looking, and it is indistinguishable to the customer from the reply sent after looking.
Check outward, in this order
- The app's own health endpoint. Does it answer, and what commit and dependencies does it report? A health endpoint that reports its database as unreachable ends the investigation here.
- Error rate and latency for the last hour in SigNoz, Grafana or Datadog, split by route if you can. An outage for a third of requests looks like health on an average.
- The edge. In Cloudflare, check the zone's analytics for the same window and the status of any recent configuration change. A DNS record, a firewall rule or a cache setting changed today is the likeliest cause of "down for me, fine for you".
- Their path specifically. If the report names a region, a browser or an endpoint, look at that slice rather than the aggregate. If it names an account, look for that account's requests in the logs by id and read the actual responses they got.
- Your dependencies. The auth provider, the payment processor, the model provider. One status page check each, and only the ones the failing path touches.
Reply with what you found
State what you checked, what it showed, and what you are doing. If it is up and they are not, say what specifically you can see succeeding — "your account's last twelve requests all answered 200, most recent two minutes ago" — because that is what turns a denial into a next step.
Stop and ask a human
- Before telling a customer nothing is wrong when you have found no positive evidence, only an absence of alerts.
- Before quoting internal error text, account ids or infrastructure names to somebody outside the company.