Best for
- RAG systems built on the default 500-tokens-and-hope
- Document sets that mix contracts, manuals, tickets and slides
- Retrieval that finds the right document but the wrong part of it
What you give it
- Your document corpus (or a representative sample) and the questions users really ask it
What you get back
- The strategy per document class: split boundaries, sizes, overlap — derived from the documents' real structure, not a universal number
- The metadata attachment plan: what each chunk carries (source, section path, version, dates) so retrieval can filter and cite
- The proof: retrieval quality on your real queries, before versus after — the strategy validated, not asserted
How it works
- Audits the corpus first: what document types exist, what structure they carry (sections, clauses, threads), and where their meaning boundaries actually fall.
- Splits along meaning, not token counts: the unit a human would quote is the unit retrieval should return — with sizes fitted per class and the trade-offs (precision versus context) made explicitly.
- Prepends location context to every chunk: the heading path, the document identity, the effective dates — an orphaned paragraph retrieves worse and cites worse than a located one.
- Proves against real queries: a labelled query set measures right-passage retrieval before and after, because chunking choices asserted without measurement are folklore.
Example
You: Our policy-and-manual RAG retrieves the right document but answers from the wrong section. Fix the chunking.
Result: The corpus audit found three document classes needing three treatments: policies (clause-aligned splits — the old fixed-size chunks cut clauses mid-sentence, which was the wrong-section bug: a chunk's first half matched, its second half answered), manuals (section-aligned with the heading path prepended to every chunk — 'Section 4.2 > Refunds > Exceptions' turns an orphan paragraph into a located one), and tickets (kept whole — splitting conversations had severed questions from their answers). Overlap kept only where structure genuinely bleeds. Proved on the 60-query set: right-passage retrieval went from 61% to 89%, and the wrong-section answers went to near zero.
Limits — please read
- Chunking is half the retrieval story; embedding and query handling are the other half (a RAG architecture pass covers the whole).
- Tables, figures and heavily-formatted content need their own treatment; the audit flags them rather than mangling them silently.
- Corpus drift (new document types) erodes fitted strategies; the class detection belongs in the ingestion path.