Task completion rate
Task completion rate is the share of jobs where an agent actually reaches the finish line with the correct end result. You give it 100 tasks, count how many it truly nailed, and that percentage is your number. The hard part is defining "done" honestly, because an agent that returns a confident answer is not the same as one that returned the right answer.
Why it matters
An agent can call every tool perfectly and still fail the user. Book the wrong flight, or write code that runs but does the wrong thing, and every step looks fine while the outcome is wrong. Completion rate scores the one thing users actually feel: did the job get done? Without it you can ship an agent that demos beautifully and quietly fails a third of the time in production.
How it works
Build a set of real tasks, each with a clear success check: a final-state assertion, an expected output, or a rubric an LLM judge can apply. Run the agent on all of them and mark each pass or fail on the outcome, ignoring whether it looked busy. Report the pass rate and slice it by task type so you can see where it breaks. Pair this with trajectory evaluation, which grades how the agent got there. Completion rate grades only whether it arrived.
A support agent is asked to cancel a subscription and issue a prorated refund. It finds the account, calls the cancel tool, and replies "all set." But it skipped the refund, so the task check fails even though the transcript reads smoothly. That one counts as zero, and your completion rate just told you the truth the demo hid.