The AI evaluation & quality glossary
Plain-English definitions of the vocabulary behind evaluating, testing, and shipping AI you can trust. Every term is a real page with a concrete example, linked to the pillar it belongs to. 52 terms and counting.
14 terms
Evaluation
Golden set · Golden dataset
A fixed set of example inputs with known-good answers you grade every new version against.
LLM-as-a-judge
Using a strong model to grade another model’s answers against a rubric, at a scale humans can’t.
Eval harness
The plumbing that runs your test cases through a model and scores the answers automatically.
Semantic similarity
Scoring how close two pieces of text are in meaning, not in exact wording.
F1 score
A single number that balances catching everything against not crying wolf.
Rubric-based grading
Grading answers against explicit written criteria instead of a gut-feel score.
Pairwise comparison
Asking “which of these two answers is better?” instead of scoring each in isolation.
Elo rating
Ranking models the way chess ranks players: by who beats whom, head to head.
Perplexity
A measure of how surprised a language model is by a piece of text.
Synthetic data generation
Using a model to invent realistic test cases when you don’t have enough real ones.
Eval dataset versioning
Treating your test set like code: tracked, reviewed, and tied to a version.
Eval-driven development · EDD
Writing the eval before the feature, so “done” means “passes the eval”.
Prompt regression testing
Re-running old cases after a prompt change to make sure nothing quietly broke.
Model card
A short spec sheet describing what a model is for, how it was tested, and its limits.
11 terms
RAG & Retrieval
RAG · Retrieval-augmented generation
Fetching relevant documents first, then letting the model answer using them.
Faithfulness · Groundedness
Whether an answer is actually supported by the sources it was given.
Context precision
Of the chunks you retrieved, how many were actually relevant.
Context recall
Whether the chunk that held the answer was retrieved at all.
Answer relevancy
Whether the answer actually addresses the question that was asked.
Chunking strategy
How you slice documents into pieces small enough to retrieve and feed to a model.
Hybrid search
Combining keyword search with meaning-based search to retrieve better context.
Reranker
A second pass that re-sorts retrieved chunks so the most useful ones come first.
Query rewriting
Rephrasing a user’s messy question into one that retrieves better.
Agentic RAG
RAG that can decide to search again, search differently, or use a tool before answering.
Embedding drift
When your search index slowly falls out of sync with the model that reads it.
9 terms
Safety
Hallucination
When a model states something false with complete confidence.
Prompt injection
Hidden instructions in user input or documents that hijack what the model does.
Jailbreak
Tricking a model into ignoring its own safety rules.
Red teaming
Deliberately attacking your own AI system to find how it fails before someone else does.
Adversarial testing
Testing with inputs designed to break the model, not the happy path.
Guardrails
Checks around a model that block bad inputs or outputs before they cause harm.
Toxicity scoring
Automatically rating output for hostility, slurs, or other harmful content.
Bias evaluation
Checking whether a model treats similar people or groups differently.
PII leakage
When a model repeats personal data it should have kept to itself.
6 terms
Agents
Agent trajectory evaluation
Grading the whole path an agent took, not just its final answer.
Tool-call accuracy
Whether an agent calls the right tool with the right arguments.
Function-calling evaluation
Testing that a model produces valid, correct calls to your functions.
Task completion rate
How often an agent actually finishes the job it was given.
Multi-turn evaluation
Testing a conversation over many back-and-forth turns, not one reply.
Structured output validation
Checking that a model’s JSON (or other format) is valid and matches your schema.
7 terms
CI / CD
CI quality gate · Eval gate
An automated check that blocks a release if quality drops below a threshold.
Ratchet
A gate that only lets quality go up: today’s score becomes tomorrow’s floor.
Flake · Flaky test
A test that passes and fails on the same code, for no real reason.
Golden-run evidence
The saved, verifiable record of a test run you can show instead of just claiming it passed.
Canary prompt
A known input you run constantly to catch a regression the moment it appears.
Shadow deployment
Running a new version alongside the old one on real traffic, without users seeing it.
Champion-challenger
Keeping the current best model in charge until a challenger proves it’s better.
5 terms
Operations
Human-in-the-loop · HITL
Putting a person at the point where an AI decision needs a human to approve it.
Model drift
When a model’s accuracy quietly decays because the world changed, not the model.
Non-determinism
Why the same prompt can give you a different answer every time you run it.
Temperature
The dial that controls how random or predictable a model’s output is.
LLM gateway · Model router
A single doorway in front of many models that handles routing, limits, and logging.
Prefer the applied version?
The Learn library turns these terms into guides, and a free mini-eval turns them into findings on your feature.