How Often Should You Re-Run Your LLM Evals
Run evals too rarely and regressions slip through; too often and the bill and the noise pile up.
Evals cost money and attention. Run them on every commit and you burn API budget while flaky tests drown the team in noise. Run them once a quarter and a bad prompt tweak ships silently, then rots for weeks before anyone notices. There is no single right cadence. You want a few eval sets, each firing on the trigger that matches what it catches.
#Split the suite before you split the schedule
The cadence question is really a suite question. You do not have one eval set. You have three, and they want different clocks.
A small smoke set of ten to thirty cases covers the paths you cannot break: the core task, the failure modes you already know, the prompt injection that burned you last time. It has to be cheap and fast enough to run on every change. A larger regression set of a few hundred cases is where subtle quality drift shows up, and it is too slow to gate every commit. A release set is the full thing, including the slow judge-graded and human-reviewed cases you only look at when something real is about to ship.
Put everything in one bucket and you are forced to pick one cadence for all of it. Every choice you make will be wrong for some of the cases.
#Match each trigger to what actually changed
Run the smoke set when a prompt, a model version, a retrieval config, or a tool definition changes. Not on every commit. A CSS edit did not touch the model, so grading fifteen answers again tells you nothing and costs you a dollar and two minutes. Gate on the inputs to the model, not on git activity.
Run the regression set nightly against whatever is on main. Nightly decouples cost from commit frequency. Ten commits or a hundred, you pay for the eval once, and you get a dated trend line instead of a wall of per-PR noise. When the nightly dips, you have a narrow window of commits to blame.
Run the release set before a release. This is the one place the slow, human-in-the-loop grading earns its keep, because a regression that reaches production costs far more than the eval that would have caught it.
#The hidden trigger everyone forgets
Your code is not the only thing that changes. The provider ships a silent update, or you bump from one snapshot to the next, and your prompts behave differently against inputs you never touched.
So a model version change is a first-class eval trigger, same as a prompt edit. Pin your model to a dated snapshot so upgrades are a decision you make on purpose, then run the regression set the moment you change the pin. Teams that float on latest get regressions with no commit to point at, which is the worst kind to debug.
Keep the nightly running even in a quiet week. A green nightly on a week you shipped nothing is not wasted. It is the control that tells you where the drift came from when it finally shows up.
#The bottom line
Cadence is a routing problem. Tie each eval set to the event that can actually break it, and the cost and noise fall out of the design instead of forcing a bad compromise. Get the smoke set fast and the nightly honest, and pre-release stops being the scramble where you find out what broke three weeks ago.