Function-calling evaluation
Function-calling evaluation checks whether a model turns a request into a call your code can actually run: the right function name, the right arguments, in the shape you expect. Here is the catch. A call that looks right and a call that runs right are two different bars, and models clear the first far more often than the second.
Why it matters
Pick the wrong tool, invent a parameter, or pass a string where you need a number, and your code either throws or quietly does the wrong thing. In an agent that chains steps, one bad call poisons everything after it. Skip this evaluation and you learn about the failure in production, when a user's action silently does nothing or fires the wrong operation.
How it works
Pair a set of user inputs with the call you expect, then score whether the model called at all when it should have, picked the correct function, and got the arguments right. Argument checks range from exact match on required fields to type and enum validation against your schema. Also test the empty case: when no tool fits, the model should not force one. Run it on every prompt or model change and watch the pass rate.
A support bot has a function issueRefund(orderId, amount). A user says "refund my last order for the full 49 dollars." A good call passes the real order ID and 49 as a number. A failing one passes "forty-nine dollars" as a string, or calls lookupOrder instead, and your refund code chokes before a cent moves.