Agents / Skills

Retrieval Quality Check

Skill

Measures whether your RAG system's retrieval actually finds the right passages — separately from whether the answers read well — because a fluent answer from the wrong source is the failure mode that erodes trust silently.

Best for

  • RAG debugging that currently starts (and ends) at the final answer
  • Knowing whether to fix retrieval or generation — they need different surgeons
  • The regression check before changing embeddings, chunking or indexes

What you give it

  • The RAG system's retrieval layer and a set of real questions (it helps label the should-retrieve passages)

What you get back

  • The scorecard: per query type, did the right passages come back, ranked high enough to be used — hit rates and ranking quality, separated
  • The failure taxonomy: the misses classified — vocabulary gaps, wrong-granularity chunks, filter failures, ranking inversions — each pointing at its own fix
  • The regression harness: the labelled set and scoring that runs on every retrieval-layer change thereafter

How it works

  1. Builds the labelled set: real queries each mapped to the passages that SHOULD be retrieved — the ground truth that makes retrieval measurable at all.
  2. Scores retrieval in isolation: hit rate at usable depth and ranking quality, sliced by query type — the answer layer stays out of this measurement entirely.
  3. Classifies every miss by mechanism: vocabulary mismatch, granularity, metadata filters, ranking, index staleness — because each class has a different fix, and the taxonomy is the repair plan.
  4. Leaves the harness installed: the set, the scoring, the per-type report — so every future retrieval change shows its delta before shipping.

Example

You: Our document assistant gives confident answers that are sometimes from the wrong policy version. Check retrieval.

Result: The check, on 80 labelled queries: overall right-passage-in-top-5 was 71%, but the taxonomy told the real story — version failures were not ranking problems at all (the version filter was simply not applied on one query path: found by the check, fixed in an hour), vocabulary-gap misses clustered on queries using customer words for policy concepts ('money back' never matching 'refund' — query expansion added), and one chunk-granularity class (long clauses split mid-condition) accounted for most ranking inversions. After the three fixes: 93%, and the wrong-version answers went to zero. The harness now gates every embedding and chunking change.

Limits — please read

  • Labelling should-retrieve passages takes expert time; it starts narrow (the highest-stakes query types) and grows.
  • Retrieval excellence does not guarantee grounded answers; the generation layer gets its own check (a grounding pass).
  • Scores reflect the labelled queries; novel query shapes enter the set through the failure-intake path.