Diagram — the LLM eval metric cheat sheetSVG · code-native
CORRECTNESS Exact match / F1 closed, deterministic answers Semantic similarity open answers where meaning matters LLM-as-judge (rubric) nuanced quality, graded at scale Pairwise / Elo ranking two prompt or model versions RETRIEVAL · RAG Faithfulness answer must be grounded in sources Answer relevancy answer actually addresses the question Context precision retrieved chunks are on-topic Context recall the needed chunk was retrieved at all SAFETY Toxicity / bias user-facing, brand-risk surfaces PII leakage the model can see sensitive data Injection resistance untrusted text enters the prompt Jailbreak rate the model has tools or authority OPERATIONAL Latency p95 it sits in a user-facing path Cost per eval / run the suite runs on every commit Output-variance same input, different answers Task completion the system takes multi-step actions
Pick the smallest set that answers "is this good enough?" for your feature — most teams need two or three of these, not all sixteen.
via Jason Teixeira — LLM Evaluation Metrics Explained