Passing Tests Is Not Proof: What Done Means for an AI Feature
Green tests mean the code ran, not that the feature is good. For AI, done needs a higher bar.
A green test suite tells you the code did what you told it to do. It says nothing about whether the feature is good. For normal software those two things stay close, because the behavior is deterministic and you can pin it down with assertions. For an AI feature they split apart, because the output is a probability distribution over language and "correct" is a range instead of a value. Ship on green alone and you have proven the plumbing works while the actual product quality goes untested.
#What a passing test actually proves
A unit test asserts that a specific input produces a specific output. That works when the answer is fixed. add(2, 2) returns 4 or it does not. The assertion is exact because the behavior is exact.
An AI feature has no fixed value. Ask a summarizer the same question twice and you get two different paragraphs, both possibly fine. So people write the only test that can pass reliably: assert the call did not throw, the response is non-empty, the JSON parses. That test goes green forever. It checks that the pipe is connected, not that anything worth drinking came out. The model could return a fluent, confident, wrong answer on every single call and your suite would stay green through all of it.
#Plumbing versus behavior
Split your checks into two buckets and keep them apart. Plumbing tests cover the deterministic shell around the model: the prompt gets built, the API gets called, retries fire, the output validates against a schema, malformed responses get caught. These should be exact assertions and they should be green. You need them.
Behavior evaluation covers what the model actually said. Is the summary accurate. Does the answer use the retrieved context. Did it refuse when it should have. This cannot be a pass/fail assertion on one run. It is a graded score across a set of cases, and it moves every time you change the prompt or swap the model. Teams get burned when a green plumbing suite stands in for behavior nobody measured. Passing the first bucket tells you nothing about the second.
#A definition of done you can act on
Done for an AI feature needs four things the test suite will not hand you. An eval set: 30 to 100 real cases with a known or graded good answer, run on every change so you watch quality move. A floor you agreed on before shipping, like "answers stay grounded in context on at least 90 percent of the retrieval set," so good enough is a number instead of a mood. A look at the failures and not just the average, because a feature at 92 percent that whiffs on the highest-value questions is not done and the mean hides it. And a plan for when the model is wrong, because it will be: a fallback, a human checkpoint on the expensive actions, or an undo.
Take a support-answer bot. Green tests prove it returns a reply and logs the ticket. Done means you graded 50 real tickets, it stayed factual on 46, the 4 misses were low-stakes, and anything touching a refund waits for a human. That is a claim you can defend. "All tests pass" is not.
#The bottom line
Keep the plumbing tests. They are real and they should stay green. Just stop letting them impersonate a quality they never measured. The honest version of done is more work. You have to build the eval set, pick the number, and read the failures nobody wants to read. Do it anyway, because the alternative is shipping on a green light that was only ever wired to the code.