JTjason.teixeira() Docs
services Book a call →
Home / Docs / The eval method / LLM Evaluation Metrics Explained
The eval method

LLM Evaluation Metrics Explained

There are dozens of LLM evaluation metrics and most of them are noise for your particular feature. This is the working guide to the ones that matter: what each measures, when to reach for it, and how they add up to a number you can gate a release on.

Every eval metric is an attempt to answer one question: is this good enough to ship, and how would I know if it quietly got worse? The metrics differ only in which kind of "good" they measure. Pick the wrong ones and you get a dashboard full of green that tells you nothing; pick the right two or three and you get a gate that actually catches a regression before your users do.

This guide sorts the useful metrics into four families, tells you when each earns its place, and shows how they roll up into the eval gate that the rest of the method is built on.

#The four families of eval metrics

Almost every metric worth running falls into one of four buckets. The cheat sheet below is the whole landscape on one screen. The right-hand line on each is the trigger: the situation where that metric stops being academic and starts being the thing standing between you and an incident.

Diagram — the LLM eval metric cheat sheetSVG · code-native
CORRECTNESS Exact match / F1 closed, deterministic answers Semantic similarity open answers where meaning matters LLM-as-judge (rubric) nuanced quality, graded at scale Pairwise / Elo ranking two prompt or model versions RETRIEVAL · RAG Faithfulness answer must be grounded in sources Answer relevancy answer actually addresses the question Context precision retrieved chunks are on-topic Context recall the needed chunk was retrieved at all SAFETY Toxicity / bias user-facing, brand-risk surfaces PII leakage the model can see sensitive data Injection resistance untrusted text enters the prompt Jailbreak rate the model has tools or authority OPERATIONAL Latency p95 it sits in a user-facing path Cost per eval / run the suite runs on every commit Output-variance same input, different answers Task completion the system takes multi-step actions
Pick the smallest set that answers "is this good enough?" for your feature — most teams need two or three of these, not all sixteen.
▸
You do not run all sixteen. A retrieval chatbot lives on faithfulness, answer relevancy, and injection resistance; a code generator lives on exact-match tests and latency. Choosing the smallest honest set is the skill.

The same landscape as a table you can copy into a planning doc:

FamilyMetricReach for it when
CorrectnessExact match / F1answers are closed and deterministic (classification, extraction)
Semantic similarityanswers are open text and meaning matters more than wording
LLM-as-judge (rubric)you need nuanced quality graded at a scale humans cannot
Pairwise / Eloyou are ranking two prompt or model versions head to head
Retrieval · RAGFaithfulness / groundednessthe answer must be supported by the retrieved sources
Answer relevancythe answer must actually address what was asked
Context precisionyou are checking whether retrieved chunks are on-topic
Context recallyou are checking whether the needed chunk was retrieved at all
SafetyToxicity / biasoutput is user-facing and carries brand risk
PII leakagethe model can see data it must never repeat
Injection resistanceuntrusted text (a document, a webpage) enters the prompt
Jailbreak ratethe model holds tools, spend authority, or private data
OperationalLatency p95it sits in a path a human is waiting on
Cost per eval / runthe suite runs on every commit and the bill compounds
Output variancethe same input can produce different answers run to run
Task completionthe system takes multi-step actions, not just single replies

#Correctness: is the answer right?

For closed tasks — a classifier, an extractor, a router — correctness is cheap and unambiguous: exact match or F1 against a labeled set, exactly like any other test. The trouble starts with open-ended output, where two very different strings can both be correct. Semantic similarity (embedding distance against a reference answer) handles paraphrase, but it rewards sounding-alike over being-right, so it is a smoke detector, not a judge. When the quality you care about is genuinely nuanced — tone, completeness, whether an explanation is actually helpful — that is the job of an LLM-as-judge, covered below.

#Retrieval: is the answer grounded?

If your feature is retrieval-augmented (a support bot, a docs assistant, anything that fetches context before answering), correctness is downstream of retrieval, and you have to measure both halves. Faithfulness asks whether the answer is actually supported by the retrieved sources or whether the model went off-script and made something up. Answer relevancy asks whether it addressed the real question. Then the two retrieval-side metrics: context precision (were the chunks you pulled on-topic?) and context recall (did you pull the chunk that actually held the answer?). A RAG system that hallucinates is almost always failing recall — the answer was never in the context to begin with. There is a deeper RAG evaluation guide for that whole pillar.

#Safety: can it be made to misbehave?

Safety metrics are the ones teams skip until an incident makes them mandatory. Toxicity and bias scoring matters the moment output is user-facing. PII leakage checks matter the moment the model can see sensitive data. And the two adversarial ones — injection resistance and jailbreak rate — matter the moment untrusted text enters the prompt or the model gains tools and authority. These are not measured with a friendly test set; they are measured with adversarial probes written specifically to make the system fail, because an attacker will.

#Operational: can you afford to run it?

A perfect eval suite you cannot afford to run on every commit is a suite that runs never. Latency (measure p95, not the average, because the average hides the failures your users feel) and cost per run decide whether the gate is sustainable. Output variance is the quietly important one: LLMs are non-deterministic, so a metric that swings five points between identical runs will either block good releases or wave bad ones through. You manage it by fixing sampling where you can and running enough samples to separate real movement from noise.

#LLM-as-judge: powerful, and quietly fallible

Using a strong model to grade another model against a rubric is the highest-leverage technique here, and it scales nuanced judgment to thousands of cases. It also has real failure modes you have to design around: judges show position bias (they favor whichever answer came first), verbosity bias (longer reads as better), and self-preference (a model rates its own family higher). Left unmanaged, a judge produces confident, consistent, wrong scores.

!
Calibrate the judge before you trust it. Grade a small human-labeled set with your judge and check agreement; randomize answer order to kill position bias; force a structured rubric with explicit criteria rather than a vibe score. An uncalibrated judge is not a metric, it is a rumor.

#Why your eval score will not match production

This is the caveat that separates people who run evals from people who trust them. Your score is only as honest as your golden set — the fixed cases you grade against. If that set was written in a quiet afternoon and production is full of typos, adversarial users, and edge cases nobody imagined, your 94% is measuring a world your users do not live in. The fix is not a cleverer metric. It is feeding real, sometimes embarrassing production cases back into the golden set on a schedule, so the thing you measure keeps drifting toward the thing your users actually do.

#How many metrics do you actually need?

Fewer than you think. Start from the failure that would hurt most — a hallucinated answer, a leaked record, a jailbreak, a latency spike — and pick the one or two metrics that catch it. Add a metric only when you can name the specific regression it exists to stop. A gate with three metrics that everyone understands beats a dashboard with twenty that nobody reads.

#From metrics to a gate

Metrics on a dashboard change nobody's behavior. Metrics wired to a threshold that blocks a merge change everything. That is the last step: pick your handful of metrics, set an honest threshold on each (usually "no worse than last known-good", a ratchet), and run them in CI so a regression stops the release instead of reaching your users. That mechanism is the CI eval gate, and it is the whole reason to measure any of this.

Want these wired into your pipeline?
Get a free mini-eval on your live AI feature — I'll show you which of these metrics your feature is missing, with real findings.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms