Eval dataset versioning
Eval dataset versioning means treating your test set like source code, tracked in version control with a full history. When your golden set lives in a file with a history, you can point to exactly which cases existed when a score was recorded. Here is what people miss: a moving test set makes every score meaningless. You can no longer tell if the model improved or the exam just got easier.
Why it matters
If someone quietly edits or deletes eval cases, your numbers drift and you have no idea why. A jump from 84 to 91 could be a real fix, or it could be three hard cases that got dropped last week. Without a version tag, you cannot reproduce an old result or trust a comparison across two runs. You end up defending a score you can no longer explain.
How it works
Keep the dataset in git next to your code, so every change shows up in a diff and goes through review. Tag or hash the exact set used for each run, and store that version next to the score. When you add cases from production, commit them with a message saying where they came from. Two runs are only comparable if they cite the same dataset version.
A refund bot scores 90 on Monday and 82 on Friday, and the team panics about a regression. The git history on the eval file shows someone added eight new edge cases Thursday afternoon. The model held steady; the exam got harder. The version tag let them prove it in two minutes instead of two days.