Best for
- Chat-with-your-documents features that confidently miss the answer
- Deciding chunking, indexing and retrieval settings with evidence
- Knowledge bases where freshness and permissions matter
What you give it
- Your document corpus and the questions users really ask
- Rules: who may see what, and what the system must refuse to answer
What you get back
- A pipeline design: ingestion, chunking fitted to your documents, indexing, retrieval, generation
- Retrieval evaluated separately from generation — you learn which half fails
- Grounding rules: answers cite sources, unknown stays unknown, permissions enforced at retrieval
How it works
- Studies the corpus and the real questions first — chunking follows document structure, not a fixed number.
- Attaches the metadata that filtering will need: source, version, date, audience, permissions.
- Evaluates retrieval alone (did the right passages come back?) before touching generation.
- Constrains generation to retrieved content with citations, and defines the honest 'not found' behaviour.
- Designs the update path: how new and changed documents enter without stale survivors.
Example
You: Our policy-assistant answers from the wrong policy version about once a day and invents clause numbers.
Result: Diagnosis: chunks split mid-clause and no version filter at retrieval. Rebuilt: clause-aligned chunking with version and effective-date metadata, retrieval filtered to the version in force, answers constrained to cite retrieved clause ids only. On the 80-question set: wrong-version answers 11 to 0, invented clauses 6 to 0, 'not found' now honestly said 4 times where the corpus has no answer.
Limits — please read
- RAG is only as good as the corpus; gaps and contradictions in your documents get surfaced, not papered over.
- Permissions enforced at retrieval require your access model to be expressible as metadata — it will tell you if it is not.
- Some questions need reasoning across many documents; it flags where plain retrieval stops sufficing.