JTjason.teixeira() Docs
services Book a call →
Home / Docs / AI Agent Reliability / How to Test AI Systems That Give Different Answers Each Run
AI Agent Reliability

How to Test AI Systems That Give Different Answers Each Run

The same input can give a different answer every run. Here is how to test something that will not sit still.

Run the same prompt twice and you can get two different answers. Sometimes both are correct, sometimes one is wrong, and your test harness has no idea which world it is in. The instinct is to pin everything down and demand an exact match. That is the wrong move. It will either flake constantly or pass on garbage. The job is not to force the model to sit still. It is to test the parts that should be stable and measure the parts that will not be.

#Test properties, not exact strings

The reflex from normal software testing is assertEqual(output, expected). For a model that rephrases itself every run, that assertion is a coin flip. You will spend your week rerunning red builds that were never real failures.

Instead, assert on properties: things that must be true of any acceptable answer, whatever the wording. If you ask for JSON, the property is that it parses and has the right keys. If you ask for a refund amount, the property is that the number matches the order, whatever the sentence around it says. If you ask for a summary, the property is that it names the people in the source and invents no new ones.

Concretely: instead of checking that a support classifier returns the string "billing_issue", check that the returned label is one of your known categories and is the expected one. The category has to be right. The prose around it is free to move.

#Separate the flaky test from the flaky model

When a run fails, you have two suspects, and people almost always blame the wrong one. Either the model produced a genuinely bad answer, or your test was too strict and a fine answer tripped it. Telling these apart is most of the work.

A good rule: if a passing answer and a failing answer are both things you would happily ship, your test is broken. Loosen the assertion. If the failing answer is one you would be embarrassed by, the test caught something real, even if it only catches it one run in ten.

That one-in-ten case is the trap. An intermittent real failure looks exactly like a flaky test from a single run. The only way to see the difference is to stop trusting single runs.

#Run it many times and measure the rate

A single pass or fail on a non-deterministic system is close to meaningless. The signal is the pass rate across many runs. Run the same case twenty or fifty times and record how often the property held. Now you have a number you can reason about: this prompt passes 96 percent of the time, that one passes 70 percent.

This changes what a test even is. Your gate is no longer "did it pass" but "did the pass rate stay above the line." You set the threshold per case based on what a wrong answer costs. A tone check might be fine at 85 percent. A check that the model never leaks another customer's data has to sit at 100 percent, and if it does not, the feature does not ship.

The honest cost: this is slower and it costs more, because you are paying for fifty runs instead of one. There is no trick around that. Sampling once to save money just means you learn about the 5 percent failure from a user instead of from your suite.

▸
Test the properties that must hold rather than the exact text, and judge each case by its pass rate over many runs rather than a single green or red. A non-deterministic system needs a statistical gate.

#The bottom line

You cannot make a model deterministic without lobotomizing the thing that made it useful, so stop trying. Pin down what must be true, let the wording drift, and measure how often reality meets your bar. It costs more than a normal test suite and it will not give you the comfort of a clean pass or fail. What it gives you instead is a real number for how often your system is right, which is the only thing worth knowing before you ship it.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms