Shift-Left Testing for AI: Catch Regressions Before Merge
The earlier a regression fails, the cheaper it is. For AI features, that means an eval on every pull request.
A regression that ships to production costs you a rollback and a bad afternoon. The same regression caught on a pull request costs one review comment. That gap is the whole argument for shift-left, and it hits AI features harder than normal code, because AI failures are quiet. A prompt tweak that looks fine can drop accuracy on a whole class of inputs, and nothing crashes to tell you. So run an eval on every pull request, the same way you run unit tests.
#Why AI regressions hide until it's too late
Normal code fails loudly. A null pointer throws. A type error blocks the build. A broken test goes red. AI features fail on a spectrum. The model still returns a fluent, confident answer. It is just wrong more often than it was last week, and only on certain inputs.
The failure is statistical. You do not catch it from one manual test, because your one test probably still passes. You catch it three weeks later when a customer forwards a screenshot, and by then you have shipped four more prompt changes and cannot tell which one did it. A pre-merge eval turns that invisible drift into a number attached to a specific commit.
#What a PR eval actually looks like
You do not need a research harness. You need a fixed set of inputs with known-good expectations, a way to run the current build against them, and a threshold that blocks the merge when the score drops.
Start small. Take thirty to fifty real examples that matter: the support tickets that must route correctly, the extractions that must pull the right order number, the answers that must cite a real policy. Write down what a correct output looks like for each. On every PR, run the feature against that set in CI and compare. If accuracy on your routing set falls from 94 percent to 88 percent, the check fails and the author sees it in the diff.
Grading is the part people overthink. For exact outputs, a string match or regex is fine and fast. For fuzzy outputs, use an LLM judge, but only after you have confirmed it agrees with your own grading on that set. An uncalibrated judge turns a real regression into a green check, which is worse than no check.
#Keep it fast, keep it honest
A gate people wait five minutes for is a gate people route around. Keep the PR eval small enough to finish in the time a normal test suite takes. Run the fifty-case smoke set on every PR. Save the thousand-case full sweep for nightly or pre-release, where a slow run is fine.
Here is the honest tradeoff. A fifty-case set will miss things. Its job is to catch the obvious cliffs, the change that tanks a whole category, before merge. It will not catch a subtle two-point regression on an edge case you never wrote down. That is the deal. You trade completeness for speed, and a signal on every change beats a perfect signal on none.
One rule keeps it real. When a regression escapes to production, add that case to the eval set before you fix it. The set grows toward the failures you actually have.
#The bottom line
Shift-left for AI is the old discipline applied to a component that fails quietly. Start with fifty cases you actually care about, a threshold that blocks, and a grader you have calibrated. Grow the set every time something slips through. It will not catch everything, and it is not supposed to. It moves the cheapest failures to the cheapest place to catch them.