The AI Incident Post-Mortem, With a Real Example
When an AI feature fails a user, a good post-mortem turns the embarrassment into a permanent test case.
Most AI features fail quietly. A user asks for something reasonable, the model gives back something wrong, and nobody hears about it until a support ticket or a screenshot on social media. The post-mortem is where you stop being embarrassed and start being useful. The whole point is to walk out with a permanent test case, so the exact failure can never ship again unnoticed.
#A worked example
Here is a real shape of incident. A customer-support assistant is asked, "Can I get a refund on my annual plan?" The user is three weeks past the 30-day refund window. The model, trying to be helpful, replies: "Yes, I've started your refund, you'll see it in 5-7 days." It has no ability to start a refund and no knowledge of the policy window. The user waits a week, no refund arrives, and now you have an angry customer plus a promise you have to honor or explain away.
Notice what actually broke. The model was fluent and confident, which is exactly why nobody caught it in the demo. It was a plausible answer that happened to be wrong, on an input the happy-path testing never included. The happy-path testing used a customer inside the refund window.
#The template, filled in
A good AI post-mortem answers five questions in plain language. What did the user ask. The exact input, copied verbatim. Here: "Can I get a refund on my annual plan?" from an out-of-window account. What did the system do. The literal output, plus what it triggered downstream. It promised a refund it could not issue.
What should it have done. Checked the account's refund eligibility, and if outside the window, said so and offered to escalate to a human. Why did it do the wrong thing. This is the honest part. The model had no tool to check eligibility and no instruction to refuse promises about actions it cannot take, so it filled the gap with a confident guess. What stops it recurring. This is the deliverable, and it is the next section.
#Turn the finding into a test
The tempting fix is to add a line to the system prompt: "Never promise refunds you cannot process." Do that, sure, but do not stop there. A prompt edit is invisible. Six weeks later someone refactors the prompt, drops that line, and the bug walks back in with nobody the wiser.
The durable fix is a test case pinned to the exact failing input. Add "refund request from an out-of-window account" to an eval set, with a check that the answer does not promise a refund and does mention the policy or an escalation. Now the failure has a tripwire. Every future model swap, prompt rewrite, and retrieval change runs against it.
One caveat: an LLM feature is non-deterministic, so a single passing run does not prove the fix holds. Run the case several times, or grade it with a checked judge, and treat a flaky pass as a fail. A test that only passes sometimes is telling you the fix is not real yet.
#The bottom line
The embarrassing incidents are the good ones. A real user just handed you an input your test set was missing. Write the five answers honestly, especially the why, then convert the finding into a pinned, repeatable test before you close the ticket. Do that every time and your eval set slowly becomes a map of every way your feature has actually hurt someone, which is the only test suite worth trusting.