Tool-call accuracy
Tool-call accuracy asks a simple question: when your agent decides to act, does it pick the right tool and fill in the right arguments? An agent can have perfect access to a refund API and still call the wrong one, or call the right one with a broken order ID. The hard part is that both failures look like normal activity until you check the actual call.
Why it matters
An agent that talks well but acts wrong is worse than useless, because it does real things in real systems. The wrong tool means it looks up weather when the user asked about shipping. Wrong arguments mean it processes a $500 refund as $5000, or cancels order 1234 instead of 1243. These are actions with consequences, and you will not catch them by reading the reply.
How it works
You build cases where you already know the correct call, then compare what the agent actually invoked against it. Two things get scored: did it choose the right tool, and did the arguments match. Argument checks range from exact matching for IDs and enums to looser checks for free-text fields. This is close to function-calling evaluation, and it usually runs on the raw tool-call payload the model emits, before any tool actually executes.
A support bot gets "cancel my most recent order." The right move is cancel_order with the order ID from the user's latest purchase. Instead it calls refund_order with yesterday's order. Tool-call accuracy flags both the wrong tool and the wrong argument, so you catch it in testing instead of after a customer's wrong order vanishes.