JTjason.teixeira() Docs
services Book a call →
Home / Docs / The eval method / RAG Evaluation Metrics Explained
The eval method

RAG Evaluation Metrics Explained

A retrieval-augmented answer can be wrong in two very different ways. These four metrics tell you which one you are looking at, so you fix the real problem instead of guessing.

RAG is two systems wearing one coat. First a retriever fetches context, then a generator writes an answer from it. A bad answer comes from one side or the other, and the fix is completely different depending on which. Grade only the final answer and you are guessing. The four metrics below split the pipeline along its real seam so you can see where it actually broke.

Diagram — where a RAG pipeline is measuredSVG · code-native
QUESTION RETRIEVE GENERATE ANSWER RETRIEVAL METRICS context precision · context recall GENERATION METRICS faithfulness · answer relevancy
A wrong answer is almost always one of two bugs: retrieval never found the right context, or generation ignored it. The metrics split cleanly along that line.

Two of the metrics watch retrieval, two watch generation. Read them as a pair of gauges: if the generation side looks fine but answers are still wrong, your problem is upstream in retrieval, and no amount of prompt tuning will save you.

#Faithfulness: did the answer stick to the sources?

Faithfulness (sometimes called groundedness) asks whether every claim in the answer is actually supported by the retrieved context. An unfaithful answer adds facts from nowhere, which is a hallucination in a nicer outfit. You score it by splitting the answer into individual claims and checking each one against the context, usually with an LLM judge or a natural-language-inference model. Low faithfulness means the generator is improvising instead of using what it was given.

#Answer relevancy: did it actually answer the question?

Answer relevancy checks the other half of a good answer: that it addresses what was asked. An answer can be perfectly faithful to the sources and still miss the point, rambling about something adjacent. This one catches the on-topic-but-useless reply. Together, faithfulness and answer relevancy pin down whether the generator did its job with the context it had.

#Context precision: was the retrieved context on-topic?

Context precision looks at the chunks the retriever pulled and asks how many were actually relevant. Low precision means the retriever is dragging in noise, which crowds the useful chunk out of the context window and distracts the generator. If precision is low, look at your chunking, your embeddings, or whether you need a reranker to push the good chunks to the top.

#Context recall: was the needed chunk retrieved at all?

Context recall is the one people skip, and it is often the real culprit. It asks whether the chunk that holds the answer made it into the retrieved set in the first place. If recall is low, the answer was never in front of the generator, so the model either says it does not know or invents something. A RAG system that hallucinates is usually failing recall. You cannot prompt your way out of a chunk that was never retrieved.

#Reading the four together to find the bug

The point of measuring all four is diagnosis. Walk them in order. If context recall is low, fix retrieval first, because nothing downstream can be right without the source. If recall is fine but context precision is low, your retriever finds the answer but buries it in noise, so rerank or tighten chunks. If retrieval looks healthy but faithfulness is low, the generator is ignoring good context, which is a prompting or model problem. And if everything is grounded but answer relevancy is low, you are answering a question next to the one that was asked. Four gauges, one clear next move.

!
Most of these metrics are themselves LLM-graded, so treat the scores as a strong signal and spot-check them against your own judgment. A metric you never calibrated is a rumor, not a measurement.

#From metrics to a gate

Scores on a dashboard change nothing. The payoff is wiring your two or three most important RAG metrics into a CI eval gate with a threshold, so a change that quietly drops faithfulness or recall fails the build instead of reaching users. For the broader picture of which metrics exist across all of LLM evaluation, see the metrics cornerstone, and for the full walkthrough there is a dedicated RAG evaluation guide.

Not sure why your RAG answers are wrong?
Get a free mini-eval on your live retrieval pipeline — I'll show you whether it is a retrieval or a generation problem, with real numbers.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms