JTjason.teixeira() Docs
services Book a call →
Home / Docs / RAG & Retrieval Quality / Why Your RAG Pipeline Hallucinates and How to Fix It
RAG & Retrieval Quality

Why Your RAG Pipeline Hallucinates and How to Fix It

A RAG system that makes things up is usually failing retrieval, not generation. Here is how to tell and how to fix it.

When a RAG system makes something up, the instinct is to blame the model. Usually that is the wrong suspect. A model handed the right passage rarely invents an answer. Far more often the retriever handed it nothing useful, and a helpful model filled the gap with a confident guess. The fix is almost never a better prompt. It is fixing what the model was given to read.

#First, prove where the failure is

Before you touch anything, run the failing question and print the exact chunks the retriever returned. Read them yourself. Ask one thing: is the answer actually in here?

This splits every hallucination into two piles. If the correct information is not in the retrieved chunks, it is a retrieval failure, and no model change will save you. If the correct information is right there and the model still contradicted it, that is a generation failure, and now a prompt or model change is worth trying. Most teams skip this step and spend a week tuning the generator for a problem that lived in the retriever. Ten minutes of reading chunks tells you which half you are in.

#The common retrieval failures

Chunking that splits the answer. You cut documents on a fixed token count, and the sentence that defines a term lands in chunk 4 while the value lands in chunk 5. Retrieval grabs one, not both, and the model sees half a fact. Chunk on structure instead, like headings and paragraphs, and add a little overlap so a thought is not sliced down the middle.

The embedding missed the match. Pure vector search fails when the user's words do not sit near the document's words in embedding space. That is exactly what happens with product codes, error strings, names, and acronyms. Someone searches ERR_4021 and semantic similarity shrugs. Add keyword search back in and blend the two. Hybrid retrieval is not fancy. It just stops you from losing exact-match terms that embeddings smear together.

The right chunk was there but ranked ninth. If you only pass the top three, it never reaches the model. Pull thirty, rerank, keep the best five. A reranker over a wider candidate set fixes more real hallucinations than most prompt work.

#The generation failures worth fixing

Sometimes retrieval did its job and the model still went off. The most fixable cause is that you never told it what to do when the context is thin. Left unsaid, a model treats "I could not find it" as failure and produces a plausible answer instead. Say it plainly in the prompt: answer only from the provided context, and if the answer is not there, say you do not know.

The other one is drowning. Stuff twenty chunks in and the real answer sits in the middle, where models reliably pay the least attention. Fewer, better-ranked chunks beat a big pile. This is why the reranker earns its place twice. It lifts retrieval quality, and it shrinks what the generator has to wade through.

▸
A RAG hallucination is a diagnosis before it is a fix. Read the retrieved chunks first, because if the answer was not in them, no amount of prompt engineering will make the model stop guessing.

#The bottom line

RAG hallucination is mostly a retrieval problem wearing a generation costume. Instrument the retriever so you can always see what it returned, and the two failure modes stop blurring together. Then you spend your effort where the fault actually is instead of tuning the half that was working fine.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms