Chunking Strategies for RAG, Compared
How you slice documents into chunks quietly decides whether retrieval works. Here are the options and the tradeoffs.
Everyone tunes the embedding model and the prompt. Almost nobody thinks hard about how the document got cut into pieces, and that decision quietly caps how good retrieval can ever be. If the answer is split across two chunks, or buried in a chunk full of unrelated text, no reranker saves you. Chunking is the cheap lever most teams ignore.
#Fixed-size chunking is the default, and it is blunt
The baseline everyone starts with is to cut every N tokens with some overlap. Split at 512 tokens, carry 50 tokens from the previous chunk into the next so a sentence that spans the boundary is not lost. It is fast, predictable, and works on any document.
The problem is it does not care what it cuts. A fixed splitter will slice through the middle of a table, end a chunk halfway through a definition, or bolt the last line of one section onto the first line of the next. Retrieval then pulls a chunk where the key sentence is a fragment, and the model fills the gap with a guess.
Overlap is the band-aid, and it mostly works. Bump it to 10 to 20 percent and most boundary casualties get repaired, because the important sentence now lives whole in at least one chunk. You pay for it with a larger index and more near-duplicate hits to dedupe.
#Sentence and structure chunking respects the document
Instead of counting tokens, split on real boundaries: sentences, paragraphs, markdown headings, list items. Pack whole sentences into a chunk until you hit a size budget, then start a new chunk on a clean break. No sentence gets cut in half.
This is the quiet sweet spot for most projects. It costs almost nothing extra and kills the worst fixed-size failures. If your source is markdown or HTML, split on headers and keep the header text inside the chunk. A chunk that starts with ## Refund Policy retrieves far better for a refund question than the same paragraph with no heading.
The catch is it needs structure to exploit. A scanned PDF flattened into one wall of text, or a transcript with no punctuation, gives the splitter nothing to grab, and you fall back to something close to fixed-size anyway.
#Semantic chunking sounds right and often is not worth it
The idea: embed each sentence, walk the document, and start a new chunk when the meaning shifts, measured by a drop in similarity between adjacent sentences. In theory each chunk is one coherent topic. In practice it is finicky.
You now have a similarity threshold to tune, and it is document-dependent. Set it too tight and you get lots of tiny chunks. Too loose and it behaves like paragraph splitting with extra steps. You also pay to embed every sentence at ingest, and the knob drifts as your corpus changes.
Be honest about the payoff. On messy documents that jump topics, like meeting notes or long threads, semantic splitting can beat sentence packing. On clean structured docs it usually ties paragraph chunking while costing more to build and maintain. Reach for it after you have measured sentence chunking, watched it fail on specific documents, and can point at why.
#The bottom line
There is no universal best chunk size, so stop hunting for one. Build a small retrieval eval: thirty real questions paired with the passages that should answer them, and measure whether your chunks surface those passages. Change one thing at a time and watch the number move. The right chunking is whatever wins on your documents, and you only know that after you measure.