JTjason.teixeira() Docs
services Book a call →
Home / Docs / AI Testing in CI/CD / Canary Releases for AI Models: A Practical Playbook
AI Testing in CI/CD

Canary Releases for AI Models: A Practical Playbook

Ship a new model to one percent of traffic before all of it. Here is how to run a canary for AI.

A new model looks better on your eval set, so you swap it in for everyone. Then support tickets spike on a use case your evals never covered. Canary releasing is the fix. Send the new model a sliver of real traffic first, watch it against the old one, and only widen the door if the numbers hold. Your offline evals are a lab. Production is the weather.

#What a canary actually is here

A canary runs two models at once on live traffic. The old model keeps serving most requests. The new one serves a small slice, one percent to start. Same inputs, real users, side by side. You are not asking whether the new model is good in the abstract. You are asking whether it is worse than the thing already running, on the traffic you actually get.

The part people skip is that routing has to be sticky per user. If a user bounces between the old and new model mid-conversation, you get incoherent behavior and you cannot attribute anything to either one. Hash the user ID into a bucket and keep them there for the whole session. One percent of users, not one percent of calls.

#Pick metrics before you ship, split by speed

AI canaries are harder than backend canaries because the thing you care about most has no instant signal. A deploy that returns 500s shows up in seconds. A model that gives subtly worse answers shows up in next week's churn.

So track two tiers. Fast guardrails you can read in minutes: error rate, p95 latency, cost per request, output length, refusal rate, and schema-valid rate if you parse the output. These catch the loud failures cheap. Then slow quality signals: thumbs-up rate, human review on a sample, an LLM judge comparing canary and control answers, and downstream conversion.

Here is the rule that works. Block automatically when any fast guardrail breaches a threshold. Hold the canary at its current percentage until you have enough slow-signal volume to trust it. Do not promote on green guardrails alone. Fast metrics being fine only means the model did not crash. It says nothing about whether it got smarter.

#Ramp on a schedule, make rollback boring

Move in steps with dwell time between them: 1 percent, 5, 25, 50, 100. Sit at each stage long enough to collect real signal. For a low-traffic product, one percent might take a day to say anything useful, so start higher or wait longer. Ramping faster than your slowest metric can respond is gambling with extra steps.

Rollback has to be a config change, not a redeploy. Keep the model choice behind a flag so cutting the canary back to zero is instant and touches nothing else. Canaries rarely fail because detection was bad. They fail because rolling back is annoying, so someone talks themselves into waiting one more hour. Make it a single toggle and that temptation disappears.

One more thing that bites: prompt changes deserve the same canary as model swaps. A reworded system prompt fails the same quiet way a new model version does, and teams canary the model while shipping the prompt to everyone at once. Route both through the same slice.

▸
A canary answers one question: is the new model worse than the one already running, on real traffic. Wire up instant rollback and slow quality metrics before you ship, because green fast guardrails only prove the model did not crash.

#The bottom line

Canary releasing does not tell you a model is good. It tells you a change is safe to widen, which is the only question that matters at deploy time. Set your thresholds, make rollback a toggle, ramp no faster than your slowest metric can talk, and canary your prompts like you canary your models. The whole point is to find the bad release on one percent of users instead of all of them.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms