Eval Gates: The Missing CI Step for AI Features
You would never ship code without CI running your tests. Most teams ship AI features with no equivalent check at all. An eval gate is that missing step.
Every serious codebase has a gate: tests run in CI, and a red build blocks the merge. AI features usually have nothing like it. A prompt gets tweaked, it looks fine in a quick manual check, and it ships. The regression shows up later as a confused user or a support ticket. An eval gate closes that hole by running your feature against known cases on every change and failing the build when quality drops.
The idea is boring on purpose. It is a unit test for behavior that happens to be fuzzy. The value is not cleverness, it is that a machine, not a tired human on a Friday, decides whether quality is still good enough to ship.
#The anatomy of a gate
A gate is four parts. A golden set of cases with known-good answers. One or more metrics that score each answer. A threshold that says what counts as passing. And an enforcement point in CI that blocks the merge when the score falls below the line. Take any one away and it stops being a gate: no golden set and you have nothing to measure, no threshold and you have a dashboard nobody reads, no enforcement and you have a suggestion.
#The ratchet: quality only goes up
The best threshold is not a fixed number you argue about. It is a ratchet: the gate stores the last known-good score and refuses anything worse. Today's result becomes tomorrow's floor. This sidesteps the endless debate about whether 88 percent is good enough and replaces it with a simpler rule that a change may not make things worse without someone consciously deciding to lower the bar.
#Where it goes in the pipeline
A gate runs where your other checks run. On every pull request for a fast subset, so feedback is quick. On a schedule overnight for the full, more expensive suite. And optionally as a hard deploy blocker for the highest-stakes features, so a failing eval can actually stop a release. The more the model can do, the closer to the deploy the gate should sit.
#The three objections, answered honestly
People resist eval gates for three real reasons. Non-determinism: the same input can score differently run to run, so a naive gate is flaky. The fix is to fix sampling where you can and run enough samples to separate real movement from noise. Cost: a huge suite on every commit gets expensive, so run a small fast set on each change and the full set nightly. Flakiness turning into theatre: if people constantly override the gate, the cases are wrong, not the people, so fix the cases. None of these are reasons to skip the gate. They are reasons to build it with a little care.
#From concept to a working gate
This is the why and the shape. When you are ready to build one, the step-by-step version lives in the tutorial on building an eval gate in CI with Promptfoo, and the animated walkthrough of how a gate behaves is the CI eval gate guide. The whole method around it is the evaluation method.