Build a Golden Set for LLM Regression Testing
Build the fixed set of known-good cases that turns "it feels better" into a number you can trust.
A golden set is a fixed collection of inputs paired with known-good answers. You grade every new version of your feature against it, so a change becomes a number instead of a hunch. This guide builds one from scratch and keeps it honest. No special tool is required, though an eval runner makes it nicer.
#Before you start
- A working LLM feature (a prompt, a chain, or a RAG pipeline).
- A handful of real examples of what users actually ask.
- Somewhere to store text files in your repo.
#Start with real cases, not invented ones
Pull thirty to a hundred real inputs from logs, support tickets, or your own testing. Cover three buckets: the common happy path, the known failures you have already hit, and the edge cases that scare you. Real cases beat clever synthetic ones, because they are the world your users actually live in.
{"id": "cancel-plan", "input": "How do I cancel my plan?", "expect": "Explains Settings > Billing > Cancel; no invented fee"}
{"id": "refund-window", "input": "Can I get a refund?", "expect": "States the 14-day window; no promises beyond it"}
{"id": "empty-question", "input": "", "expect": "Asks a clarifying question; does not guess"}#Decide how each case is graded
Each case needs a way to score the answer. Use exact match or a keyword check for closed answers, semantic similarity for open ones, and an LLM-as-judge with a rubric for the nuanced cases. Write the grading rule down next to the case, so the standard is explicit and not living in your head.
def grade(case, answer, judge):
# closed check first: fast and free
if case.get("must_include"):
if case["must_include"].lower() not in answer.lower():
return 0.0
# nuanced check: ask a judge against the written expectation
verdict = judge(
question=case["input"],
answer=answer,
rubric=case["expect"],
)
return verdict.score # 0.0 to 1.0#Run the set and record one score
Run every case through the current version, grade each answer, and roll them into a single number like average score or pass rate. Save the per-case results too, so a drop points you at the exact cases that broke rather than just a lower number.
python golden/run.py --version prompt-v3
# -> golden set: 0.91 (was 0.88). 2 cases regressed: refund-window, empty-question#Grow it from production
The set is never finished. Every time production surprises you with a wrong or unsafe answer, add that exact case with the correct expectation. Over time the golden set drifts toward reality, and the score starts to mean something closer to how the feature behaves for real users.
#Watch out for
- A stale golden set lies. If it was written once and never updated, a high score just means you are good at last year's questions.
- Do not grade open-ended answers with exact match. You will fail correct answers for using different words and end up chasing wording instead of meaning.
- Keep the set in version control so a score is always tied to a specific version of the cases. See the tutorial on versioning golden datasets.
#What you built
You have a golden set that turns any change to your feature into a comparable score, plus a per-case breakdown that tells you what moved. It is the backbone of every other eval technique here, and it costs almost nothing to start. Wire it into a gate next, so the score actually blocks a regression.