Best for
- Incidents that keep recurring in new costumes
- Post-incident writeups that currently stop at 'human error'
- Making 'why did this happen' a method instead of a mood
What you give it
- What happened: timeline, evidence, the people available to ask
- Openness to answers that implicate process, not persons
What you get back
- The causal chain written down: trigger, enabling conditions, root causes
- Contributing factors ranked by leverage: fixing which would have prevented this?
- Actions that match causes — each owned, dated and checkable
How it works
- Builds the timeline first from evidence, not recollection — memory reorders events.
- Asks why repeatedly but follows branches: most failures have several contributing causes, not one neat chain.
- Treats 'human error' as the start of a question (why was the error easy to make and hard to catch?), never the conclusion.
- Matches each fix to the causal level it addresses, and rejects fixes that only address the trigger.
Example
You: A config change took production down for 40 minutes on Tuesday. Run the analysis.
Result: The chain: the change was valid but applied to the wrong environment (trigger), possible because environment names differ by one character and the tool shows no diff preview (enabling), and because config changes bypass review entirely (root). Three actions matched to the three levels — renaming environments, diff preview on apply, lightweight config review — not 'be more careful', which was the draft conclusion before the analysis.
Limits — please read
- Honesty in equals quality out: analysis in a blame culture gets polite fiction — it will say when it detects that.
- Some causes exceed your control (vendor internals); they get named as constraints to design around.
- The method finds causes; funding the fixes is a decision it can only frame.