JTjason.teixeira() Glossary
Home / Learn / Glossary / Chunking strategy
RAG & Retrieval

Chunking strategy

Chunking strategy is how you cut a document into pieces before storing them for retrieval. Each chunk becomes a unit your system can search for and hand to the model. The hard part is size. Too big and you bury the answer in noise. Too small and you slice it in half, so no single piece makes sense on its own.

Why it matters

Retrieval can only return chunks you already created, so bad chunking quietly caps how good your answers can ever get. Split a policy down the middle and the model gets the setup without the exception, or the number without the condition attached to it. You see this as answers that are technically sourced but still wrong, and no amount of prompt tuning fixes it, because the right text was never in one place.

How it works

You pick a size, often a few hundred tokens, plus an overlap so a sentence cut at the edge still shows up whole in the next chunk. Smarter versions split on real boundaries like paragraphs, headings, or code functions instead of a blind character count. Measure it with context recall (did the needed text get retrieved at all) and context precision (was it near the top), then adjust size and boundaries until both climb.

In practice

A refund bot keeps saying refunds take 5 to 7 days. The real doc reads "5 to 7 days for cards, 30 days for bank transfers," but the chunker split it right after "5 to 7 days." Now the bank-transfer line lives in a different piece and rarely gets retrieved with the first one. Chunk on the full policy paragraph instead, and the model finally sees both cases at once.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy