Best for
- Error trackers with 4,000 unread events nobody opens anymore
- Finding the real signal after a rough deploy
- Turning log noise into a short, ranked fix list
What you give it
- The logs or error-tracker export, and the window you care about
- What 'bad' means for you: which flows and users matter most
What you get back
- Errors clustered by underlying cause — 4,000 events become 12 distinct problems
- Each cluster ranked by real impact: how many users, which flows, growing or stable
- Onset analysis per cluster: when it started and what shipped or changed right before
How it works
- Normalises and clusters events by root signature — same cause, one cluster, regardless of surface variation.
- Separates volume from impact: a loud harmless error must not outrank a quiet one that blocks checkout.
- Builds the timeline per cluster and correlates onsets with deploys, config changes and traffic shifts.
- Distinguishes new, worsening, stable and already-fixed — each gets different urgency.
- Reports a ranked fix list with the evidence trail, ready to hand to whoever fixes it.
Example
You: Since Tuesday the error tracker is on fire. Here is the export. What is actually broken?
Result: 11 clusters. One is 78% of the volume but harmless (a retried timeout that succeeds). Three matter: a payment-callback failure (240 users, started 14 minutes after Tuesday's deploy), a null crash on profiles imported from the old system, and a slow burn growing 8% daily that predates Tuesday entirely. Ranked list with evidence, onset times and suspect changes.
Limits — please read
- It triages from what is logged; invisible failures (missing logging) get named as blind spots, not guessed.
- Correlation with a deploy is a strong lead, not a conviction — it says which it has.
- Fixing the top of the list is a separate task (a good companion: a root-cause debugging agent).