How Reliable Should an AI Agent Actually Be
The right reliability bar for an agent depends entirely on the blast radius of its worst action.
Most reliability debates about AI agents start in the wrong place. They argue about the model's accuracy in a vacuum. The number that matters is not how often the agent is right. It is what happens the one time it is wrong. An agent that drafts internal notes and an agent that moves money can run the exact same model and need wildly different reliability bars, because a mistake costs different amounts. Set the bar by the blast radius, then work backward to what you actually need.
#Blast radius is the real spec
Blast radius is the worst thing that happens when the agent takes its worst action on a bad day. Not the average case. The tail. Ask a plain question. If this agent does the wrong thing on a real input nobody tested, who notices, how much does it cost, and how fast can you undo it.
A support agent that tags tickets has a tiny blast radius. A wrong tag is annoying and a human fixes it in seconds. Point the same architecture at issuing refunds and the blast radius is large, because a wrong refund is real money leaving the account and clawing it back is awkward at best. Same model, same prompt quality, completely different requirement. What sets the bar is the action, and how smart the model is barely enters into it.
#Reversibility beats accuracy
Here is the part teams skip. An action that is cheap to undo needs far less reliability than one that is permanent. A 95% accurate agent that writes draft emails a human sends is fine, because the human is the undo button. A 95% accurate agent that sends those emails directly is not, because 1 in 20 wrong emails is now in a customer's inbox and you cannot recall it.
So before you chase another accuracy point, ask if you can make the action reversible instead. Soft deletes over hard deletes. Drafts over sends. A staged change with a rollback in front of a live write. Reversibility is usually cheaper to engineer than reliability, and it caps the blast radius no matter how the model fails.
#Set the bar with a small grid
You do not need a framework. You need two axes. How reversible the action is, and how costly the worst outcome is. Sort each agent action into that grid before you build.
Low cost and reversible, like tagging or summarizing, can run fully automatic at ordinary model reliability. High cost or irreversible, like payments, deletions, or anything a customer sees unedited, needs a human checkpoint or a hard limit in code. A better prompt will not save you here. The messy middle, medium cost and reversible, is where you spend your evaluation effort, because that is where the tradeoff is real.
A rule you can apply today. Any action the agent cannot undo without a human should require a human to approve it, until you have measured its failure rate on inputs you did not cherry-pick. Irreversible plus unmeasured equals gated.
#The bottom line
There is no universal reliability number for AI agents, and anyone quoting one is selling something. Score each action by cost and reversibility. Automate the cheap and undoable ones, and put a person or a hard limit in front of the ones you cannot take back. Do that and a mediocre model becomes safe to ship, while an ungated irreversible action stays dangerous no matter how good the model gets.