JTjason.teixeira() Docs
services Book a call →
Home / Docs / AI Testing in CI/CD / Shift-Left Testing for AI: Catch Regressions Before Merge
AI Testing in CI/CD

Shift-Left Testing for AI: Catch Regressions Before Merge

The earlier a regression fails, the cheaper it is. For AI features, that means an eval on every pull request.

A regression that ships to production costs you a rollback and a bad afternoon. The same regression caught on a pull request costs one review comment. That gap is the whole argument for shift-left, and it hits AI features harder than normal code, because AI failures are quiet. A prompt tweak that looks fine can drop accuracy on a whole class of inputs, and nothing crashes to tell you. So run an eval on every pull request, the same way you run unit tests.

#Why AI regressions hide until it's too late

Normal code fails loudly. A null pointer throws. A type error blocks the build. A broken test goes red. AI features fail on a spectrum. The model still returns a fluent, confident answer. It is just wrong more often than it was last week, and only on certain inputs.

The failure is statistical. You do not catch it from one manual test, because your one test probably still passes. You catch it three weeks later when a customer forwards a screenshot, and by then you have shipped four more prompt changes and cannot tell which one did it. A pre-merge eval turns that invisible drift into a number attached to a specific commit.

#What a PR eval actually looks like

You do not need a research harness. You need a fixed set of inputs with known-good expectations, a way to run the current build against them, and a threshold that blocks the merge when the score drops.

Start small. Take thirty to fifty real examples that matter: the support tickets that must route correctly, the extractions that must pull the right order number, the answers that must cite a real policy. Write down what a correct output looks like for each. On every PR, run the feature against that set in CI and compare. If accuracy on your routing set falls from 94 percent to 88 percent, the check fails and the author sees it in the diff.

Grading is the part people overthink. For exact outputs, a string match or regex is fine and fast. For fuzzy outputs, use an LLM judge, but only after you have confirmed it agrees with your own grading on that set. An uncalibrated judge turns a real regression into a green check, which is worse than no check.

#Keep it fast, keep it honest

A gate people wait five minutes for is a gate people route around. Keep the PR eval small enough to finish in the time a normal test suite takes. Run the fifty-case smoke set on every PR. Save the thousand-case full sweep for nightly or pre-release, where a slow run is fine.

Here is the honest tradeoff. A fifty-case set will miss things. Its job is to catch the obvious cliffs, the change that tanks a whole category, before merge. It will not catch a subtle two-point regression on an edge case you never wrote down. That is the deal. You trade completeness for speed, and a signal on every change beats a perfect signal on none.

One rule keeps it real. When a regression escapes to production, add that case to the eval set before you fix it. The set grows toward the failures you actually have.

▸
Treat an AI eval like a test suite. A small, fast, known-answer set runs on every pull request and blocks the merge when the score drops. Catching the regression in the diff is cheap. Catching it from a customer screenshot is not.

#The bottom line

Shift-left for AI is the old discipline applied to a component that fails quietly. Start with fifty cases you actually care about, a threshold that blocks, and a grader you have calibrated. Grow the set every time something slips through. It will not catch everything, and it is not supposed to. It moves the cheapest failures to the cheapest place to catch them.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms