JTjason.teixeira() Glossary
Home / Learn / Glossary / Agentic RAG
RAG & Retrieval

Agentic RAG

Agentic RAG is retrieval-augmented generation where the model decides what to do before it answers, instead of running one fixed search and hoping it worked. It can search again with different words, pull from a second source, or call a tool when the first grab comes back thin. Here's the catch: the deciding part is the model judging its own retrieval, so it can also judge wrong and loop or stall.

Why it matters

Plain RAG does one lookup. If that lookup misses, the model answers anyway from whatever weak scraps it got, and you get a confident wrong answer. Real questions often need more than one hop: find the account, then its plan, then that plan's rules. Agentic RAG lets the system notice I don't have enough yet and go get the rest. That's the difference between handling hard questions and only handling easy ones.

How it works

A loop wraps the retriever. The model reads the question, retrieves, then checks whether it has enough to answer. If not, it rewrites the query, picks a different source or tool, and retrieves again, with a hard cap on iterations so it can't spin forever. Measure it on the parts, not only the final text. Did retrieval improve on each hop (context recall and precision)? Was the answer grounded in what it found (faithfulness)? Watch cost and latency too, since every extra hop is another model call.

In practice

A support bot gets "why was I charged twice in March?" One search for "double charge" returns the generic billing FAQ, which doesn't answer it. An agentic version notices the gap, looks up the user's actual March transactions, spots a failed retry that posted twice, and explains that specific event instead of reciting policy.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy