Best for
- Teams with no on-call playbook and occasional chaos
- Solo operators who need a second brain at 3am
- Turning incidents from panic into procedure
What you give it
- The symptom and access to logs, dashboards and the system
- Decisions where stabilising costs something (rollback, feature off, degraded mode)
What you get back
- A stabilisation move proposed in minutes, with its trade-off stated
- A running incident log: timeline, actions, evidence, decisions
- Status updates ready to send, and a structured handoff to the post-incident review
How it works
- Establishes the blast radius first: who is affected, how badly, since when.
- Separates stabilisation from diagnosis — stop the bleeding, then find the cause.
- Proposes the lowest-risk mitigating move and states what it costs.
- Logs every action, observation and decision with timestamps as it goes.
- Keeps communication on a cadence so stakeholders stop interrupting the fixers.
Example
You: Checkout is failing for about a third of customers. It started 20 minutes ago.
Result: Correlated the start time with a dependency's latency spike; proposed serving the cached catalogue and queueing orders (stabilised in 12 minutes); kept the timeline; two status updates drafted; the full evidence log handed to the review with the root cause confirmed.
Limits — please read
- It commands process and analysis; actions on production run through your access and your approval gates.
- Mitigation trade-offs (lose a feature vs lose uptime) are business calls — proposed, never assumed.
- It is as good as the observability you have; blind spots get named for the review.