Agents / Skills

Runbook Writer

Skill

Writes operational runbooks that work at 3am: for each alarm or task, what it means, what to check in order, the exact commands, and when to escalate — tested against the system so the night shift is not debugging the document.

Best for

  • Alerts whose response procedure lives in one veteran's memory
  • On-call rotations where every incident is improvisation
  • Routine operations (restore, rotate, scale) done rarely enough to be forgotten

What you give it

  • The procedures as they exist — in heads, chats or stale wikis — and access to verify commands

What you get back

  • One runbook per alarm or task: meaning, impact, checks in order, actions with exact commands, escalation line
  • Commands verified against the real system — copy-paste ready, placeholders made obvious
  • The decision points pre-decided: when to restart, when to roll back, when to wake the senior

How it works

  1. Writes per trigger: the alarm or task is the entry point, because that is how the reader arrives.
  2. Orders checks by diagnostic yield: the first three commands should split the likely causes.
  3. Verifies every command against the real system — a runbook with broken commands is sabotage with a table of contents.
  4. Pre-decides the judgement calls: thresholds, safe actions per finding, and the explicit escalation line with who and how.

Example

You: Write runbooks for our top six alerts; currently the answer is 'call the person who left'.

Result: Six runbooks, each tested in dry-run: the queue-depth alarm's book (what it means, the three checks that distinguish producer-flood from consumer-death — each a pasted command, the two safe actions per diagnosis, and the depth threshold that means wake the lead), plus five more. The next queue incident: resolved in 11 minutes by the engineer hired last month.

Limits — please read

  • Runbooks encode known failure modes; the novel incident still needs a thinking human — the book says when it has run out.
  • Destructive actions carry their warnings and approval gates inline; it will not write casual data-loss commands.
  • Drift is real: it dates each book and wires the review trigger (every use, or quarterly).