JTjason.teixeira() Docs
services Book a call →
Home / Docs / AI Testing in CI/CD / Setting SLOs for AI Features: What Good Enough Means
AI Testing in CI/CD

Setting SLOs for AI Features: What Good Enough Means

You cannot manage what you will not define. SLOs turn "the AI feels off" into a number the team agrees on.

"The AI feels off lately" is not a bug report. It is a vibe, and vibes do not survive a standup. An SLO is the fix. You set a target in advance for how the feature behaves, write it as a number, and agree on it before anyone gets defensive. It does not make the AI better. It makes disagreement about the AI resolvable.

#An SLO is a promise, not a hope

A service-level objective is a target for a measurable property of your feature, plus the threshold that counts as good enough. "95% of support replies are factually correct" is an SLO. "The AI should be accurate" is a wish. One has a number and a denominator. The other starts an argument.

The hard part is not picking the number. It is picking what to measure. Latency, error rate, and uptime are easy and you should track them for AI too. But the property that actually matters is quality, and quality does not ship with a built-in metric. You have to define it, usually with a graded eval set and often an LLM judge you have calibrated against humans. If you cannot measure the thing, you do not have an SLO. You have a slogan.

#Set the target from the human baseline, not from 100%

The instinct is to aim for perfect and feel bad about the gap. That is the wrong anchor. Ask instead how well the current process does, and what a wrong answer actually costs.

A human support agent is not 100% accurate. If your team gets it right 92% of the time, an AI at 95% is an upgrade, and demanding 99.9% invents a standard you never held people to. A feature that drafts a legal clause and a feature that suggests a playlist have very different costs of being wrong, so they deserve different targets even at the same accuracy. Put the SLO where the cost of failure and the value of the feature balance. Write the reasoning next to the number. In three months someone will ask why it is 95 and not 98, and "it felt right" is not an answer that holds.

#Pick a few SLOs, and give each an error budget

You cannot have an SLO on everything. Pick the two or three properties that would make you pull the feature if they broke. For a RAG assistant that is usually correctness, groundedness (did it stick to the retrieved docs), and a safety or refusal rate. Latency and cost sit alongside as their own targets. Fifteen SLOs is zero SLOs, because nobody knows which one blocks a release.

Then use the gap as a budget. A 95% correctness target means a 5% error budget. You have explicitly decided 5% wrong answers is tolerable. That is not an admission of failure. It is permission to ship. Inside budget, you ship features. Burn through it, and work stops until quality is fixed. This is the part teams skip, and it is what turns an SLO from a poster on the wall into a decision rule.

▸
An SLO is a measurable property, a target set from the real cost of being wrong, and an error budget that decides when you ship and when you stop. Without the number and the budget, you are still shipping on feel.

#The bottom line

SLOs will not tell you why the model regressed or how to fix it. They tell you, before the argument starts, whether it regressed at all and whether that is allowed. Set two or three, tie each to a cost you can name, and let the error budget make the call. It turns "the AI feels off" into either "we are inside budget, ship it" or "we are over, stop and fix it." A team can act on both of those.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms