JTjason.teixeira() Glossary
Home / Learn / Glossary / RAG
RAG & Retrieval

RAG

also called Retrieval-augmented generation

RAG means you fetch relevant documents first, then hand them to the model and ask it to answer using them. Instead of relying on whatever the model memorized in training, you give it the exact source text at question time. The catch: the answer is only as good as what retrieval pulls. A bad fetch quietly poisons a reply that still sounds fluent.

Why it matters

A raw model answers from frozen memory that never saw your private data. Ask it about your refund policy or last week's changelog and it will guess, often confidently wrong. RAG grounds the answer in your actual documents, so it quotes your real policy instead of inventing one. You update its behavior by editing the documents, with no retraining.

How it works

Two stages. First you index your documents: split them into chunks and store embeddings. Then at query time you retrieve the closest chunks, often blending vector and keyword matching with hybrid search. Those chunks go into the prompt as context, and the model answers from them. You measure the pipeline in halves: retrieval quality with context precision and context recall, generation quality with faithfulness and answer relevancy.

In practice

A support bot gets asked "how long do I have to return a jacket?" It searches the help center, pulls the returns article that says 14 days, and drops that text into the prompt. The model answers "14 days" straight from the retrieved doc. If the right article never surfaces, the fix is a retrieval bug, and no amount of prompt tweaking will save it.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy