Five things, or you're guessing.
A pass/fail RAG evaluation harness has to score all five of these on a fixed question set — ideally a golden set with known-correct sources — and gate the release on the results. Skip one and you've left a whole class of failure untested.
- 01Citation coverage. What fraction of answer claims carry a citation that actually supports them? Not "has a footnote" — the cited chunk must contain the claim. Target: every load-bearing sentence is traceable to a retrieved source.
- 02Retrieval precision & recall. Did the retriever surface the chunks that contain the answer (recall), and how much noise came with them (precision)? A perfect generator can't save a retriever that never fetched the right passage.
- 03Groundedness / faithfulness. Is every statement entailed by the retrieved context, or did the model add "knowledge" from pre-training that isn't in the sources? This is where hallucination hides even when citations are present.
- 04Abstain-when-no-evidence. When the corpus genuinely doesn't contain the answer, does the system say so — or does it confabulate? Test with out-of-corpus questions and require an explicit "insufficient evidence" response.
- 05Latency & cost per query. Retrieval + rerank + generation each add milliseconds and tokens. Measure p50/p95 latency and cost per query, because an accurate system nobody can afford to run is still a failed system.