JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Build a Golden Set for LLM Regression Testing
How-to guides

Build a Golden Set for LLM Regression Testing

Build the fixed set of known-good cases that turns "it feels better" into a number you can trust.

A golden set is a fixed collection of inputs paired with known-good answers. You grade every new version of your feature against it, so a change becomes a number instead of a hunch. This guide builds one from scratch and keeps it honest. No special tool is required, though an eval runner makes it nicer.

#Before you start

  • A working LLM feature (a prompt, a chain, or a RAG pipeline).
  • A handful of real examples of what users actually ask.
  • Somewhere to store text files in your repo.

#Start with real cases, not invented ones

Pull thirty to a hundred real inputs from logs, support tickets, or your own testing. Cover three buckets: the common happy path, the known failures you have already hit, and the edge cases that scare you. Real cases beat clever synthetic ones, because they are the world your users actually live in.

golden/cases.jsonl
{"id": "cancel-plan", "input": "How do I cancel my plan?", "expect": "Explains Settings > Billing > Cancel; no invented fee"}
{"id": "refund-window", "input": "Can I get a refund?", "expect": "States the 14-day window; no promises beyond it"}
{"id": "empty-question", "input": "", "expect": "Asks a clarifying question; does not guess"}

#Decide how each case is graded

Each case needs a way to score the answer. Use exact match or a keyword check for closed answers, semantic similarity for open ones, and an LLM-as-judge with a rubric for the nuanced cases. Write the grading rule down next to the case, so the standard is explicit and not living in your head.

golden/grade.py
def grade(case, answer, judge):
    # closed check first: fast and free
    if case.get("must_include"):
        if case["must_include"].lower() not in answer.lower():
            return 0.0
    # nuanced check: ask a judge against the written expectation
    verdict = judge(
        question=case["input"],
        answer=answer,
        rubric=case["expect"],
    )
    return verdict.score  # 0.0 to 1.0

#Run the set and record one score

Run every case through the current version, grade each answer, and roll them into a single number like average score or pass rate. Save the per-case results too, so a drop points you at the exact cases that broke rather than just a lower number.

terminal
python golden/run.py --version prompt-v3
# -> golden set: 0.91 (was 0.88). 2 cases regressed: refund-window, empty-question

#Grow it from production

The set is never finished. Every time production surprises you with a wrong or unsafe answer, add that exact case with the correct expectation. Over time the golden set drifts toward reality, and the score starts to mean something closer to how the feature behaves for real users.

#Watch out for

  • A stale golden set lies. If it was written once and never updated, a high score just means you are good at last year's questions.
  • Do not grade open-ended answers with exact match. You will fail correct answers for using different words and end up chasing wording instead of meaning.
  • Keep the set in version control so a score is always tied to a specific version of the cases. See the tutorial on versioning golden datasets.

#What you built

You have a golden set that turns any change to your feature into a comparable score, plus a per-case breakdown that tells you what moved. It is the backbone of every other eval technique here, and it costs almost nothing to start. Wire it into a gate next, so the score actually blocks a regression.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms