Every RAG answer starts as a chunk. The retriever never sees your documents as you do. It sees only the pieces you cut them into, each represented by one vector. If a fact is split across two pieces, or diluted inside a piece that covers ten topics, no amount of clever search later will recover it cleanly. Chunking is the first decision in the pipeline, and the hardest one to undo.
Why split at all?
Two forces push towards smaller pieces.
Precision of search. One embedding summarises one piece of text. If an entire 40-page policy becomes a single vector, that vector represents an average of leave rules, travel rules, the code of conduct and payroll dates all at once. A question about notice periods matches that average weakly, and it matches every other large document weakly too. Smaller chunks each say one thing, so a focused question can find a focused match.
Budget of the prompt. Whatever you retrieve goes into the model's context. Sending entire documents for every question wastes tokens, adds cost and latency, and buries the useful sentence in noise.
One force pushes the other way: context. A chunk that is too small ("It is 60 days.") is precise but meaningless on its own, because the reader cannot tell what "it" refers to. Chunking is the search for pieces small enough to be specific and large enough to make sense on their own.
Fixed-size chunking with overlap
The simplest method cuts every N characters (or tokens): characters 0–1,000, then 1,000–2,000, and so on. It is fast and predictable, and it ignores meaning entirely, so it will happily cut a sentence in half.
Overlap limits the damage. Instead of starting each chunk where the last one ended, you step back a little. With a chunk size of 1,000 and an overlap of 200, the chunks cover 0–1,000, then 800–1,800, then 1,600–2,600, and so on. A sentence that straddles a boundary now appears whole in at least one chunk. The cost is duplication: here about 25% more text to embed and store (1,000 characters for every 800 new ones), and near-duplicate chunks that sometimes both appear in the top results.
Recursive splitting: respect natural boundaries
Recursive splitting keeps the size limit but prefers to cut at a natural boundary. It tries a list of separators in order of preference:
- a blank line (paragraph break),
- a single newline,
- a sentence end or a space,
- and only as a last resort, a raw character position.
If a paragraph fits within the limit, it becomes a chunk. If it doesn't, the splitter falls back to the next separator within that paragraph, and so on recursively. The result stays close to the target size while preserving paragraphs and sentences wherever possible. This is why recursive splitting is the sensible default for prose, and a classic interview question asks for the difference: a plain character splitter cuts exactly at the size, while a recursive splitter cuts at the best boundary near the size.
Structure-aware splitting: let the document tell you
Many documents already declare their structure. Markdown has headings, HTML has tags, policies have numbered sections, and contracts have clauses. A structure-aware splitter cuts along those lines, so each chunk is one section, carrying its heading as context ("Working Hours", "Code of Conduct").
A robust production pattern combines the two:
- Split by section (heading or clause).
- Within a section, keep each paragraph as a chunk if it fits under a cap (say 1,200 characters).
- If a paragraph exceeds the cap, split it recursively.
- Attach the section title and document name to every chunk as metadata, and often prepend them to the text before embedding, so that "It is 60 days" becomes "Notice period: It is 60 days".
The cap matters because one oversized paragraph that mixes several ideas embeds into a blurred vector and matches nothing well.
Semantic chunking: split where the topic changes
The most adaptive method ignores formatting and looks at meaning directly. Embed each sentence, then walk through the document comparing each sentence with the next. While consecutive sentences are similar, they belong to the same chunk. Where similarity drops sharply (a common rule is a drop larger than the 95th percentile of all drops), the topic has shifted, and a new chunk starts.
Semantic chunking produces coherent, single-topic chunks even from badly formatted text such as transcripts or scraped pages. It costs an embedding call per sentence at indexing time, and its chunk sizes vary widely, so it is usually combined with a maximum size.
Choosing a chunk size
There is no universal right size, but there are good starting points and a way to tune them.
- Start around 300–1,000 tokens with 10–20% overlap for prose. FAQs and short policies prefer the small end, and narrative documents (reports, papers) the larger end.
- Match the size of the answer. If typical answers are one or two sentences, large chunks waste context. If answers need a whole procedure, small chunks scatter it.
- Respect the embedding model's limit. Text beyond the model's maximum input length is silently truncated and never embedded.
- Tune with evaluation, not intuition. Module 5 gives you retrieval metrics such as recall@k. Rebuild the index at two or three chunk sizes and keep the one that scores best on your own questions.
Retrieval and reading want different sizes. A popular pattern is to embed small chunks for precise matching but return their parent section (or a window of neighbouring chunks) to the model. You search with precision and read with context. Keep a parent ID in each chunk's metadata to make this possible.
Cut the same handbook four ways and check whether the answer survives in one piece.