Best for
- Before optimising anything — to find out what is actually slow
- Settling the 'I bet it is the database' argument with data
- Learning where a request's 800ms actually lives
What you give it
- The slow path, a way to run it realistically, and your profiling tools (or it works with what the stack has built in)
What you get back
- The time budget decomposed: where the milliseconds live, by layer and by function
- The honest distinction: dominant costs versus noise — and whether the bottleneck is even in your code
- A findings list ranked by potential win, each with the evidence line that proves it
How it works
- Reproduces the slowness under realistic conditions first — profiles of toy runs produce toy conclusions.
- Captures at the right layers: wall-clock decomposition across the request, then CPU/allocation/query detail where the time concentrates.
- Reads the profile with discipline: inclusive vs exclusive time, call counts, the difference between hot and merely wide.
- Verifies the headline finding by a targeted second measurement before anyone optimises anything.
Example
You: The order-summary endpoint takes 900ms and everyone has a theory. Profile it.
Result: The decomposition: 610ms in 14 sequential database round-trips (one per line item — the classic), 140ms serialising a field the response does not include, 90ms in the framework, 60ms actual logic. Nobody's theory survived: the suspected 'slow query' was 11ms. Findings ranked: batch the lookups (est. 500ms), drop the dead field (140ms) — with the profile traces attached.
Limits — please read
- Profiling observes; fixing is the next task (a performance-tuning agent pairs well).
- Instrumentation overhead can distort; it cross-checks with low-overhead methods where precision matters.
- Intermittent slowness needs capture at the bad moment; the session includes the trap-setting for that.