LLM Evaluation Metrics Explained
There are dozens of LLM evaluation metrics and most of them are noise for your particular feature. This is the working guide to the ones that matter: what each measures, when to reach for it, and how they add up to a number you can gate a release on.
Every eval metric is an attempt to answer one question: is this good enough to ship, and how would I know if it quietly got worse? The metrics differ only in which kind of "good" they measure. Pick the wrong ones and you get a dashboard full of green that tells you nothing; pick the right two or three and you get a gate that actually catches a regression before your users do.
This guide sorts the useful metrics into four families, tells you when each earns its place, and shows how they roll up into the eval gate that the rest of the method is built on.
#The four families of eval metrics
Almost every metric worth running falls into one of four buckets. The cheat sheet below is the whole landscape on one screen. The right-hand line on each is the trigger: the situation where that metric stops being academic and starts being the thing standing between you and an incident.
The same landscape as a table you can copy into a planning doc:
| Family | Metric | Reach for it when |
|---|---|---|
| Correctness | Exact match / F1 | answers are closed and deterministic (classification, extraction) |
| Semantic similarity | answers are open text and meaning matters more than wording | |
| LLM-as-judge (rubric) | you need nuanced quality graded at a scale humans cannot | |
| Pairwise / Elo | you are ranking two prompt or model versions head to head | |
| Retrieval · RAG | Faithfulness / groundedness | the answer must be supported by the retrieved sources |
| Answer relevancy | the answer must actually address what was asked | |
| Context precision | you are checking whether retrieved chunks are on-topic | |
| Context recall | you are checking whether the needed chunk was retrieved at all | |
| Safety | Toxicity / bias | output is user-facing and carries brand risk |
| PII leakage | the model can see data it must never repeat | |
| Injection resistance | untrusted text (a document, a webpage) enters the prompt | |
| Jailbreak rate | the model holds tools, spend authority, or private data | |
| Operational | Latency p95 | it sits in a path a human is waiting on |
| Cost per eval / run | the suite runs on every commit and the bill compounds | |
| Output variance | the same input can produce different answers run to run | |
| Task completion | the system takes multi-step actions, not just single replies |
#Correctness: is the answer right?
For closed tasks — a classifier, an extractor, a router — correctness is cheap and unambiguous: exact match or F1 against a labeled set, exactly like any other test. The trouble starts with open-ended output, where two very different strings can both be correct. Semantic similarity (embedding distance against a reference answer) handles paraphrase, but it rewards sounding-alike over being-right, so it is a smoke detector, not a judge. When the quality you care about is genuinely nuanced — tone, completeness, whether an explanation is actually helpful — that is the job of an LLM-as-judge, covered below.
#Retrieval: is the answer grounded?
If your feature is retrieval-augmented (a support bot, a docs assistant, anything that fetches context before answering), correctness is downstream of retrieval, and you have to measure both halves. Faithfulness asks whether the answer is actually supported by the retrieved sources or whether the model went off-script and made something up. Answer relevancy asks whether it addressed the real question. Then the two retrieval-side metrics: context precision (were the chunks you pulled on-topic?) and context recall (did you pull the chunk that actually held the answer?). A RAG system that hallucinates is almost always failing recall — the answer was never in the context to begin with. There is a deeper RAG evaluation guide for that whole pillar.
#Safety: can it be made to misbehave?
Safety metrics are the ones teams skip until an incident makes them mandatory. Toxicity and bias scoring matters the moment output is user-facing. PII leakage checks matter the moment the model can see sensitive data. And the two adversarial ones — injection resistance and jailbreak rate — matter the moment untrusted text enters the prompt or the model gains tools and authority. These are not measured with a friendly test set; they are measured with adversarial probes written specifically to make the system fail, because an attacker will.
#Operational: can you afford to run it?
A perfect eval suite you cannot afford to run on every commit is a suite that runs never. Latency (measure p95, not the average, because the average hides the failures your users feel) and cost per run decide whether the gate is sustainable. Output variance is the quietly important one: LLMs are non-deterministic, so a metric that swings five points between identical runs will either block good releases or wave bad ones through. You manage it by fixing sampling where you can and running enough samples to separate real movement from noise.
#LLM-as-judge: powerful, and quietly fallible
Using a strong model to grade another model against a rubric is the highest-leverage technique here, and it scales nuanced judgment to thousands of cases. It also has real failure modes you have to design around: judges show position bias (they favor whichever answer came first), verbosity bias (longer reads as better), and self-preference (a model rates its own family higher). Left unmanaged, a judge produces confident, consistent, wrong scores.
#Why your eval score will not match production
This is the caveat that separates people who run evals from people who trust them. Your score is only as honest as your golden set — the fixed cases you grade against. If that set was written in a quiet afternoon and production is full of typos, adversarial users, and edge cases nobody imagined, your 94% is measuring a world your users do not live in. The fix is not a cleverer metric. It is feeding real, sometimes embarrassing production cases back into the golden set on a schedule, so the thing you measure keeps drifting toward the thing your users actually do.
#How many metrics do you actually need?
Fewer than you think. Start from the failure that would hurt most — a hallucinated answer, a leaked record, a jailbreak, a latency spike — and pick the one or two metrics that catch it. Add a metric only when you can name the specific regression it exists to stop. A gate with three metrics that everyone understands beats a dashboard with twenty that nobody reads.
#From metrics to a gate
Metrics on a dashboard change nobody's behavior. Metrics wired to a threshold that blocks a merge change everything. That is the last step: pick your handful of metrics, set an honest threshold on each (usually "no worse than last known-good", a ratchet), and run them in CI so a regression stops the release instead of reaching your users. That mechanism is the CI eval gate, and it is the whole reason to measure any of this.