Infrastructure
Roll back a bad deploy
Put the last known-good build back in front of users and prove it is serving, before spending any time on why the new one broke.
By Toolspoke
Skill procedure
Rolling back is the fastest fix available and it is reversible, so it comes before diagnosis, not after it. The mistake this procedure exists to prevent is the one where a rollback is triggered, reported as done, and the old container is still serving because nothing checked.
Steps
- Name the good commit. List the application's deployments and find the most recent one that finished successfully and was live long enough to be trusted — usually the deploy before the one that broke. Write the commit SHA down before you touch anything.
- Check nothing irreversible shipped in between. Read the diff between the good commit and the current one for schema changes and data migrations. A rollback past a migration that dropped or rewrote a column does not restore the old behaviour, it produces a second outage. If you find one, stop and ask — this is the one case where forward is the only way.
- Redeploy the good commit through the platform that runs the service, and watch the deployment to completion rather than firing it and moving on.
- Prove it is serving. Call the app's health endpoint and read the commit it reports back. If the endpoint does not report a commit, compare the running container's image digest to the one you deployed. A deployment marked successful by the platform is not evidence: a container that crashes on boot leaves the previous one serving, and the platform shows green.
- Say so in one line on the incident channel: what is live now, which commit, and that the broken change is unmerged or reverted on the default branch so nobody redeploys it by accident.
Stop and ask a human
- A migration sits between the two commits (step 2).
- The rollback target is more than a day old — that is usually a sign the real problem is elsewhere and a wide rollback will cost more than it saves.
- The service has no health endpoint reporting a commit, so you cannot prove what is serving.
What done looks like
The health endpoint reports the good commit, the incident channel knows, and the broken change cannot be redeployed without somebody choosing to.