Best for
- The Saturday-night change that must not become a Sunday incident
- Making one person's plan survivable by whoever actually executes it
- Changes where 'we'll figure it out live' has burned you before
What you give it
- The change: what moves, the systems involved, the window available
What you get back
- A numbered runbook: every step with its command, expected result, and verification
- Gates and abort criteria: where to check, what 'wrong' looks like, when to pull out
- The rollback as a first-class script — plus the point of no return marked in red
How it works
- Scripts every step with its exact command, expected output and the check that proves it worked.
- Front-loads everything front-loadable: pre-flight the day before, so the window spends time only on what needs the window.
- Defines abort criteria per gate in advance — the 2am brain executes decisions, it does not make them.
- Writes the rollback with the same rigour as the forward path, and marks where it stops being available.
Example
You: We move the primary database to the new server in Saturday's 2am window. Write the runbook.
Result: A 23-step runbook: pre-flight checks Friday (backup verified restorable, replication lag zero, the new server's config diffed against old), the window script with per-step verification and timings, three gates with abort criteria ('replication lag above 30s at step 14: abort path B'), the rollback script for every stage before step 19 — and step 19 marked POINT OF NO RETURN with the final go/no-go checklist in front of it.
Limits — please read
- A runbook encodes the plan you have; rehearsing it against a copy is the step that finds what the plan missed — it will insist.
- Steps needing credentials or physical access are marked for their holders by name.
- Surprises outside the script need a human decision; the runbook says who to wake.