EN·ES·PT
Field guide · RAG evaluation

How to actually test a RAG system.

Most retrieval-augmented generation demos are graded by vibe: someone reads three answers, they look plausible, ship it. That's not evaluation — that's optimism. A RAG system fails in specific, testable ways: it cites the wrong chunk, it invents a fact the retrieved context never said, or it answers confidently when it should have said "I don't have that." Each of those failure modes has a check that catches it. This is the checklist, and it's grounded in a dashboard I built where citation coverage is 100% by construction — there is no code path that can generate an uncited sentence.

5
dimensions every RAG eval must cover
100%
citation coverage on the reference build
0
uncited generation paths (by construction)
1
live demo you can grade yourself

The 100% figure is structural, not a benchmark score: the pipeline cannot emit an answer sentence that isn't tied to a retrieved source. Verify it on the case-studies.html build.

The checklist · what a real RAG eval measures

Five things, or you're guessing.

A pass/fail RAG evaluation harness has to score all five of these on a fixed question set — ideally a golden set with known-correct sources — and gate the release on the results. Skip one and you've left a whole class of failure untested.

  • 01Citation coverage. What fraction of answer claims carry a citation that actually supports them? Not "has a footnote" — the cited chunk must contain the claim. Target: every load-bearing sentence is traceable to a retrieved source.
  • 02Retrieval precision & recall. Did the retriever surface the chunks that contain the answer (recall), and how much noise came with them (precision)? A perfect generator can't save a retriever that never fetched the right passage.
  • 03Groundedness / faithfulness. Is every statement entailed by the retrieved context, or did the model add "knowledge" from pre-training that isn't in the sources? This is where hallucination hides even when citations are present.
  • 04Abstain-when-no-evidence. When the corpus genuinely doesn't contain the answer, does the system say so — or does it confabulate? Test with out-of-corpus questions and require an explicit "insufficient evidence" response.
  • 05Latency & cost per query. Retrieval + rerank + generation each add milliseconds and tokens. Measure p50/p95 latency and cost per query, because an accurate system nobody can afford to run is still a failed system.
RAG evaluation flow A user query flows into retrieval, which returns ranked chunks. A groundedness gate then either produces a grounded, cited answer when evidence is sufficient, or abstains when it is not. Query user question Retrieval embed + search Ranked chunks rerank top-k Evidence gate Grounded + cited Abstain no evidence sufficient insufficient

The gate is the whole game. Answers only ship when retrieved evidence supports every claim; otherwise the system abstains instead of confabulating. Removing the abstain branch is the single most common cause of "confidently wrong" RAG.

The matrix · failure mode → the check that catches it

Every way RAG breaks, and its test.

RAG failure modes mapped to the evaluation check that catches each
Failure modeWhat the user seesThe check that catches it
Uncited claimA confident sentence with no source you can verify.Citation coverage
Wrong-chunk citationA footnote that points to a passage which doesn't support the claim.Citation → claim entailment
Missed retrieval"I don't know" when the answer was in the corpus all along.Retrieval recall
Noisy retrievalAnswer drifts because irrelevant chunks crowded the context window.Retrieval precision
Hallucinated factA plausible detail the sources never actually contained.Groundedness / faithfulness
Confident confabulationA full answer to a question the corpus can't support.Abstain-when-no-evidence
Stale indexYesterday's answer to today's document.Freshness / re-index check
Unaffordable accuracyCorrect answers that take 8 seconds and cost real money each.Latency & cost per query

A serious harness runs this whole matrix on a fixed golden set on every change, and fails the build when any row regresses — the same gate-on-regression discipline behind my QA work.

The proof · citation coverage by construction

100% cited, because there's no other path.

The design move

Instead of generating prose and hoping citations attach, the pipeline binds every emitted answer segment to the retrieved chunk it came from. If a claim has no supporting chunk, it doesn't get written — the system abstains on that point rather than inventing one.

Why it's structural

Citation coverage isn't a score I tuned toward — it's a property of the code path. There is no uncited generation route to regress. That's the difference between "we usually cite" and "we cannot not cite."

See it running

The research dashboard case study walks through the build. And the site's own live eval demo lets you paste an AI answer and watch it get graded against the same kind of checks in real time.

▸ see the 100%-cited build ▸ try the live eval demo coverage is structural · verify it yourself
Why take this checklist from me

I gate my own releases on evidence.

The abstain-and-cite discipline above isn't theory. My QA OS caught 15 high/critical CVEs in its own dependency set, blocked its own release, and was patched the same day to 3,759 tests passing across 13 of 13 gates. My SDET regression suite runs 37 of 37 specs, zero flakes, in 15.3 seconds. This whole site runs its own QA — 100+ checks, axe-clean — and carries a live eval demo you can break yourself.

15 CVEs caught, release blocked 3,759 tests · 13/13 gates 37/37 specs · 0 flakes · 15.3s ISTQB CT-AI ISTQB Test Automation Engineer
Related · more on testing AI

Worried your RAG system is guessing?

Book a call and I'll build you an evaluation harness that scores all five dimensions and gates your releases — or grade your AI's answer live, right now.

Book a call → try the live eval demo → see the 100%-cited build →
worked example

What the gate actually catches.

A real RAG failure isn't a crash — it's a confident, wrong answer that reads fine. Here's one the eval catches:

Query

"What's our refund window for enterprise plans?"

✕ retrieved the wrong chunk

Top match was the consumer policy ("30-day money-back"). The model answered "Yes — 30 days" — fluent, confident, and wrong for enterprise. A human skims it and ships it.

✓ gate catches it, then corrected

Faithfulness scoring flags that the answer isn't grounded in an enterprise source — fails the citation gate. Retrieval is tuned to fetch the order-form terms; the corrected answer cites the right clause.

The point: the wrong answer looked identical in quality to the right one. Only a grounding check told them apart.

frequently asked

Questions, answered.

How is testing RAG different from a plain LLM?
RAG has two failure surfaces: retrieval (did it fetch the right context?) and generation (did it answer faithfully and cite it?). You must measure both — a confident answer from the wrong chunk is still wrong.
How do you measure faithfulness and hallucination?
Each answer is scored against its retrieved context with LLM-as-judge plus deterministic citation checks, so an answer that isn’t grounded in a real source fails even when it reads well.
Can the RAG eval run in CI?
Yes. The suite gates merges — a retrieval or prompt change that drops citation coverage goes red before it ships.
What do we get at the end?
A graded RAG eval suite in your repo, a scored dashboard, a CI gate, and a runbook your team can operate.
What is the difference between citation coverage and groundedness?
Citation coverage asks whether the answer cites its sources at all; groundedness asks whether those cited sources actually support each claim. An answer can cite a chunk and still misread it — so the two are scored separately.
Which retrieval metrics do you measure — context precision and recall?
Context precision (did the retrieved chunks belong?) and context recall (did retrieval miss a chunk the answer needed?). Both run against a labeled question set, so a retriever change moves a number instead of a vibe.
Do you fix the retrieval pipeline or only measure it?
Measurement comes first, but the eval points straight at the fix — chunking, embeddings, reranking, or the prompt. I can implement those changes and re-run the suite to prove the gain.
How do you stop the model from inventing facts?
Every claim is checked against the retrieved context, and answers with no supporting source fail the grade. Paired with an abstention test, the system is scored on saying “I don’t know” rather than guessing.
Does this work with my vector database and embedding model?
Yes — the eval sits above your stack and treats retrieval as a black box, so any vector database or embedding model works. Swapping either becomes a measurable experiment rather than a blind change.
What deliverables do I get from a RAG evaluation?
A graded eval suite in your repo, a scored dashboard covering retrieval and groundedness, a CI gate that blocks regressions, and a runbook your team can operate without me.