JTjason.teixeira() Glossary
Home / Learn / Glossary / Golden set
Evaluation

Golden set

also called Golden dataset

A golden set is a fixed collection of example inputs paired with the answers you already know are good. Every time you change a prompt, swap a model, or tweak your retrieval, you run the whole set again and compare. It is the closest thing an AI feature has to a regression test.

Why it matters

Without one, you are flying on vibes. A change that fixes one case almost always breaks another, and you will not notice until a user does. A golden set turns "it feels better" into "it scored 91, up from 88, and these three cases got worse." That is the difference between guessing and knowing.

How it works

Start small and real: 30 to 100 cases that cover your common paths, your known failures, and the edge cases that scare you. Each case is an input plus a reference answer or a pass/fail rule. You grade new versions against it with exact matching, a similarity score, or an LLM judge, depending on how open-ended the answers are. The set is never finished. The best cases come from production, especially the embarrassing ones.

In practice

A support bot keeps confidently giving the wrong refund window. You add that exact question to the golden set with the correct answer attached. Now any future prompt change that reintroduces the bug fails the set immediately, before it ships, instead of after a customer complains.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy