Best for
- Alerts whose response procedure lives in one veteran's memory
- On-call rotations where every incident is improvisation
- Routine operations (restore, rotate, scale) done rarely enough to be forgotten
What you give it
- The procedures as they exist — in heads, chats or stale wikis — and access to verify commands
What you get back
- One runbook per alarm or task: meaning, impact, checks in order, actions with exact commands, escalation line
- Commands verified against the real system — copy-paste ready, placeholders made obvious
- The decision points pre-decided: when to restart, when to roll back, when to wake the senior
How it works
- Writes per trigger: the alarm or task is the entry point, because that is how the reader arrives.
- Orders checks by diagnostic yield: the first three commands should split the likely causes.
- Verifies every command against the real system — a runbook with broken commands is sabotage with a table of contents.
- Pre-decides the judgement calls: thresholds, safe actions per finding, and the explicit escalation line with who and how.
Example
You: Write runbooks for our top six alerts; currently the answer is 'call the person who left'.
Result: Six runbooks, each tested in dry-run: the queue-depth alarm's book (what it means, the three checks that distinguish producer-flood from consumer-death — each a pasted command, the two safe actions per diagnosis, and the depth threshold that means wake the lead), plus five more. The next queue incident: resolved in 11 minutes by the engineer hired last month.
Limits — please read
- Runbooks encode known failure modes; the novel incident still needs a thinking human — the book says when it has run out.
- Destructive actions carry their warnings and approval gates inline; it will not write casual data-loss commands.
- Drift is real: it dates each book and wires the review trigger (every use, or quarterly).