Best for
- Alert channels the team muted in self-defence
- Incidents that keep arriving without any alert firing
- Thresholds set once in 2021 and never questioned
What you give it
- Alert history (what fired when, what was done) and incident history (what actually went wrong)
What you get back
- The audit: every alert scored by its record — true activations, ignored firings, and the incidents it slept through
- Retuned rules: thresholds from real baselines, symptom-based where cause-based cried wolf, severities honest
- The deletions, with reasons written — removing a dead alert is a decision, not an accident
How it works
- Scores every rule by its history: when it fired, what a human did, and whether any incident occurred that it should have caught.
- Treats fired-and-ignored as the signal it is: the rule is wrong, the threshold is wrong, or the thing does not matter.
- Rebuilds the important alerts on symptoms (what users experience) with causes demoted to dashboards for diagnosis.
- Derives thresholds from measured baselines with the seasonality your system actually has.
Example
You: We get 400 alerts a week and still missed last month's outage. Fix the alerting.
Result: The audit: 400 weekly firings come from 31 rules — 6 rules produce 83% of the noise and led to action zero times this quarter (deleted, reasons recorded); 9 thresholds moved from round-number folklore to baseline-derived; the missed outage traced to cause-based alerts that all stayed green while users suffered — two symptom alerts added at the user-visible edge. Result four weeks later: 38 alerts that week, every page acted on, and the next real degradation caught in 4 minutes.
Limits — please read
- The audit needs history; thin records mean conservative retuning and better record-keeping as a finding.
- Deleting alerts takes nerve; every deletion carries its written reason so it can be argued and reversed.
- Alert hygiene decays; the audit becomes a cadence or the noise returns.