Reference
Glossary
Plain-English definitions — no jargon for its own sake. Every term is deep-linkable; hover a term and click the # to grab its anchor.
- #Golden set
- A curated collection of real inputs paired with agreed-good outputs, used as the source of truth for scoring an AI feature. see also →
- #LLM-as-judge
- Using an evaluation model to score another model’s output against criteria like faithfulness, relevance, and safety. see also →
- #Faithfulness / grounding
- Whether an answer is actually supported by its source material, rather than invented. The core RAG-quality question. see also →
- #Hallucination
- A confident, fluent answer that is not true or not grounded in any source.
- #Prompt injection
- An input crafted to make the model ignore its instructions and do something else. see also →
- #Jailbreak
- An attempt to bypass a model’s safety rules, often via role-play or a fake “no-rules” persona. see also →
- #CI quality gate
- An automated check in your pipeline that blocks a merge or deploy when a quality score drops below a floor. see also →
- #Ratchet
- A gate whose passing floor only ever moves up, so quality can’t silently erode. see also →
- #Flake
- A test that passes and fails without any code change; flaky suites train teams to ignore red. see also →
- #Golden run / evidence
- A saved, reproducible test run (traces, screenshots) that proves a result. see also →
- #Human-in-the-loop
- A design where a person approves an AI action at a defined risk point. see also →
- #RAG
- Retrieval-augmented generation: grounding answers in retrieved documents instead of model memory. see also →