Agents / Skills

RAG Chunking Strategy

Skill

Chooses how your documents get split for retrieval — by structure, size and overlap fitted to what your documents actually are — and proves the choice against real queries, because chunking is where most RAG quality is silently won or lost.

Best for

  • RAG systems built on the default 500-tokens-and-hope
  • Document sets that mix contracts, manuals, tickets and slides
  • Retrieval that finds the right document but the wrong part of it

What you give it

  • Your document corpus (or a representative sample) and the questions users really ask it

What you get back

  • The strategy per document class: split boundaries, sizes, overlap — derived from the documents' real structure, not a universal number
  • The metadata attachment plan: what each chunk carries (source, section path, version, dates) so retrieval can filter and cite
  • The proof: retrieval quality on your real queries, before versus after — the strategy validated, not asserted

How it works

  1. Audits the corpus first: what document types exist, what structure they carry (sections, clauses, threads), and where their meaning boundaries actually fall.
  2. Splits along meaning, not token counts: the unit a human would quote is the unit retrieval should return — with sizes fitted per class and the trade-offs (precision versus context) made explicitly.
  3. Prepends location context to every chunk: the heading path, the document identity, the effective dates — an orphaned paragraph retrieves worse and cites worse than a located one.
  4. Proves against real queries: a labelled query set measures right-passage retrieval before and after, because chunking choices asserted without measurement are folklore.

Example

You: Our policy-and-manual RAG retrieves the right document but answers from the wrong section. Fix the chunking.

Result: The corpus audit found three document classes needing three treatments: policies (clause-aligned splits — the old fixed-size chunks cut clauses mid-sentence, which was the wrong-section bug: a chunk's first half matched, its second half answered), manuals (section-aligned with the heading path prepended to every chunk — 'Section 4.2 > Refunds > Exceptions' turns an orphan paragraph into a located one), and tickets (kept whole — splitting conversations had severed questions from their answers). Overlap kept only where structure genuinely bleeds. Proved on the 60-query set: right-passage retrieval went from 61% to 89%, and the wrong-section answers went to near zero.

Limits — please read

  • Chunking is half the retrieval story; embedding and query handling are the other half (a RAG architecture pass covers the whole).
  • Tables, figures and heavily-formatted content need their own treatment; the audit flags them rather than mangling them silently.
  • Corpus drift (new document types) erodes fitted strategies; the class detection belongs in the ingestion path.