JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Version Golden Datasets Like Code
How-to guides

Version Golden Datasets Like Code

Treat your test set like code: tracked, reviewed, and tied to a version, so a score always means something.

You will put your eval dataset under Git, tag it, and print that tag next to every score so a number always points at a known set of cases. Once this is in place, "we scored 0.91" means something, because you can check out the exact data that produced it. This uses plain Git plus one small change to your eval runner. Budget about twenty minutes.

#Before you start

  • A golden dataset in text files (JSONL, CSV, or YAML).
  • Git installed and the dataset living in a repo.
  • An eval runner or script that reads those files and prints a score.

#Keep the dataset in text

Git works best when it can diff your data line by line. Store cases as JSONL or CSV, one record per line, so a reviewer can see exactly which case changed in a pull request. Avoid spreadsheets and pickle files, since a one-word fix shows up as an unreadable blob.

golden/cases.jsonl
{"id": "cancel-plan", "input": "How do I cancel my plan?", "expect": "Explains Settings > Billing > Cancel; no invented fee"}
{"id": "refund-window", "input": "Can I get a refund?", "expect": "States the 14-day window; no promises beyond it"}

#Commit the data as its own change

Treat a dataset edit like a code edit. Commit the cases on their own, with a message that says what moved and why. When someone asks why the score dropped last Tuesday, the commit history is the answer.

terminal
git add golden/cases.jsonl
git commit -m "golden: add refund-window case, tighten cancel expectation"

#Tag a dataset version

A tag is a stable name for one exact state of the files. Tag the dataset whenever you want a version you can refer back to, and use an annotated tag so it carries a date and message. Now golden-v3 means one specific set of cases forever, even after you keep editing.

terminal
git tag -a golden-v3 -m "golden set v3: 42 cases"
git push origin golden-v3

#Stamp the version onto every score

A score is meaningless without the dataset version beside it. Have your runner read the current Git description and print it with the result. The command below returns the nearest tag plus the short commit, so even an uncommitted run is traceable.

golden/version.py
import subprocess

def dataset_version():
    out = subprocess.run(
        ["git", "describe", "--tags", "--always", "--dirty"],
        capture_output=True, text=True, check=True,
    )
    return out.stdout.strip()  # e.g. 'golden-v3' or 'golden-v3-2-gab12cd-dirty'

if __name__ == "__main__":
    print(f"dataset: {dataset_version()}")

#Record the version with the result

Write the score and the dataset version together, so the pair travels as one record. The '-dirty' suffix is a gift here. It tells you the run used edited-but-uncommitted cases, which is exactly the run you should not trust in a report.

terminal
python golden/run.py
# -> dataset: golden-v3  score: 0.91 (pass rate 38/42)

#Watch out for

  • A '-dirty' or long-hash version means the score came from uncommitted data. That number is fine for local tinkering, but do not paste it into a report as if it were reproducible.
  • Comparing two scores across different tags is comparing two different tests. A jump from 0.88 to 0.94 can just mean you deleted the hard cases, so always diff the tags before you celebrate.
  • Git chokes on huge or binary datasets. If your set grows past a few thousand rows or includes images, use Git LFS or DVC and tag the pointer files instead of stuffing megabytes into history.

#What you built

You now have a golden dataset that lives in Git, gets tagged like a release, and prints its own version next to every score. A number is no longer a rumor. It is tied to a commit anyone can check out and rerun. Next, wire this version string into your CI eval gate, so each pull request records which dataset it was graded against.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms