JTjason.teixeira() Docs
services Book a call →
Home / Docs / The eval method / Eval Gates: The Missing CI Step for AI Features
The eval method

Eval Gates: The Missing CI Step for AI Features

You would never ship code without CI running your tests. Most teams ship AI features with no equivalent check at all. An eval gate is that missing step.

Every serious codebase has a gate: tests run in CI, and a red build blocks the merge. AI features usually have nothing like it. A prompt gets tweaked, it looks fine in a quick manual check, and it ships. The regression shows up later as a confused user or a support ticket. An eval gate closes that hole by running your feature against known cases on every change and failing the build when quality drops.

Diagram — the eval gateSVG · code-native
The eval gate AI output runs through an eval battery — correctness, safety, hallucination, and regression checks — into a CI quality gate. Green ships. Red is blocked and sent back to fix and re-run. Nothing ships on a hunch. AI OUTPUT raw, unproven EVAL BATTERY Correctness Safety Hallucination Regression CI QUALITY GATE GATE SHIP green → deploy BLOCKED red → fix & re-run
This is the differentiator. Anyone can ship an AI feature. The gate is what proves it still works after the next prompt change — and it ships only when every check is green.

The idea is boring on purpose. It is a unit test for behavior that happens to be fuzzy. The value is not cleverness, it is that a machine, not a tired human on a Friday, decides whether quality is still good enough to ship.

#The anatomy of a gate

A gate is four parts. A golden set of cases with known-good answers. One or more metrics that score each answer. A threshold that says what counts as passing. And an enforcement point in CI that blocks the merge when the score falls below the line. Take any one away and it stops being a gate: no golden set and you have nothing to measure, no threshold and you have a dashboard nobody reads, no enforcement and you have a suggestion.

#The ratchet: quality only goes up

The best threshold is not a fixed number you argue about. It is a ratchet: the gate stores the last known-good score and refuses anything worse. Today's result becomes tomorrow's floor. This sidesteps the endless debate about whether 88 percent is good enough and replaces it with a simpler rule that a change may not make things worse without someone consciously deciding to lower the bar.

#Where it goes in the pipeline

A gate runs where your other checks run. On every pull request for a fast subset, so feedback is quick. On a schedule overnight for the full, more expensive suite. And optionally as a hard deploy blocker for the highest-stakes features, so a failing eval can actually stop a release. The more the model can do, the closer to the deploy the gate should sit.

#The three objections, answered honestly

People resist eval gates for three real reasons. Non-determinism: the same input can score differently run to run, so a naive gate is flaky. The fix is to fix sampling where you can and run enough samples to separate real movement from noise. Cost: a huge suite on every commit gets expensive, so run a small fast set on each change and the full set nightly. Flakiness turning into theatre: if people constantly override the gate, the cases are wrong, not the people, so fix the cases. None of these are reasons to skip the gate. They are reasons to build it with a little care.

#From concept to a working gate

This is the why and the shape. When you are ready to build one, the step-by-step version lives in the tutorial on building an eval gate in CI with Promptfoo, and the animated walkthrough of how a gate behaves is the CI eval gate guide. The whole method around it is the evaluation method.

Want an eval gate on your feature?
Get a free mini-eval on your live AI feature, or book a call and I'll wire a gate into your pipeline.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms