Diagram — the LLM eval metric cheat sheet SVG · code-native
CORRECTNESS
Exact match / F1
closed, deterministic answers
Semantic similarity
open answers where meaning matters
LLM-as-judge (rubric)
nuanced quality, graded at scale
Pairwise / Elo
ranking two prompt or model versions
RETRIEVAL · RAG
Faithfulness
answer must be grounded in sources
Answer relevancy
answer actually addresses the question
Context precision
retrieved chunks are on-topic
Context recall
the needed chunk was retrieved at all
SAFETY
Toxicity / bias
user-facing, brand-risk surfaces
PII leakage
the model can see sensitive data
Injection resistance
untrusted text enters the prompt
Jailbreak rate
the model has tools or authority
OPERATIONAL
Latency p95
it sits in a user-facing path
Cost per eval / run
the suite runs on every commit
Output-variance
same input, different answers
Task completion
the system takes multi-step actions
Pick the smallest set that answers "is this good enough?" for your feature — most teams need two or three of these, not all sixteen.