Best for
- RAG debugging that currently starts (and ends) at the final answer
- Knowing whether to fix retrieval or generation — they need different surgeons
- The regression check before changing embeddings, chunking or indexes
What you give it
- The RAG system's retrieval layer and a set of real questions (it helps label the should-retrieve passages)
What you get back
- The scorecard: per query type, did the right passages come back, ranked high enough to be used — hit rates and ranking quality, separated
- The failure taxonomy: the misses classified — vocabulary gaps, wrong-granularity chunks, filter failures, ranking inversions — each pointing at its own fix
- The regression harness: the labelled set and scoring that runs on every retrieval-layer change thereafter
How it works
- Builds the labelled set: real queries each mapped to the passages that SHOULD be retrieved — the ground truth that makes retrieval measurable at all.
- Scores retrieval in isolation: hit rate at usable depth and ranking quality, sliced by query type — the answer layer stays out of this measurement entirely.
- Classifies every miss by mechanism: vocabulary mismatch, granularity, metadata filters, ranking, index staleness — because each class has a different fix, and the taxonomy is the repair plan.
- Leaves the harness installed: the set, the scoring, the per-type report — so every future retrieval change shows its delta before shipping.
Example
You: Our document assistant gives confident answers that are sometimes from the wrong policy version. Check retrieval.
Result: The check, on 80 labelled queries: overall right-passage-in-top-5 was 71%, but the taxonomy told the real story — version failures were not ranking problems at all (the version filter was simply not applied on one query path: found by the check, fixed in an hour), vocabulary-gap misses clustered on queries using customer words for policy concepts ('money back' never matching 'refund' — query expansion added), and one chunk-granularity class (long clauses split mid-condition) accounted for most ranking inversions. After the three fixes: 93%, and the wrong-version answers went to zero. The harness now gates every embedding and chunking change.
Limits — please read
- Labelling should-retrieve passages takes expert time; it starts narrow (the highest-stakes query types) and grows.
- Retrieval excellence does not guarantee grounded answers; the generation layer gets its own check (a grounding pass).
- Scores reflect the labelled queries; novel query shapes enter the set through the failure-intake path.