Agents / Agents

Incident Commander

Agent

Runs a production incident with discipline: stabilise first, diagnose with evidence, communicate at a steady cadence, and capture the timeline — so the outage ends sooner and teaches more.

Best for

  • Teams with no on-call playbook and occasional chaos
  • Solo operators who need a second brain at 3am
  • Turning incidents from panic into procedure

What you give it

  • The symptom and access to logs, dashboards and the system
  • Decisions where stabilising costs something (rollback, feature off, degraded mode)

What you get back

  • A stabilisation move proposed in minutes, with its trade-off stated
  • A running incident log: timeline, actions, evidence, decisions
  • Status updates ready to send, and a structured handoff to the post-incident review

How it works

  1. Establishes the blast radius first: who is affected, how badly, since when.
  2. Separates stabilisation from diagnosis — stop the bleeding, then find the cause.
  3. Proposes the lowest-risk mitigating move and states what it costs.
  4. Logs every action, observation and decision with timestamps as it goes.
  5. Keeps communication on a cadence so stakeholders stop interrupting the fixers.

Example

You: Checkout is failing for about a third of customers. It started 20 minutes ago.

Result: Correlated the start time with a dependency's latency spike; proposed serving the cached catalogue and queueing orders (stabilised in 12 minutes); kept the timeline; two status updates drafted; the full evidence log handed to the review with the root cause confirmed.

Limits — please read

  • It commands process and analysis; actions on production run through your access and your approval gates.
  • Mitigation trade-offs (lose a feature vs lose uptime) are business calls — proposed, never assumed.
  • It is as good as the observability you have; blind spots get named for the review.