Agent trajectory evaluation
Agent trajectory evaluation grades the whole path an agent took to reach an answer, every step along the way. It looks at which tools it called, in what order, what it read, and how it recovered when something broke. Here's the catch: two agents can land on the same correct answer while one got there by luck and the other actually reasoned its way there.
Why it matters
If you only check the final answer, you reward accidents. An agent that guesses right after three wrong tool calls looks identical to one that nailed it cleanly, until the guesser hits a case where luck runs out. Trajectory evaluation catches the messy middle: redundant API calls burning tokens, a wrong tool that happened to work anyway, a loop that almost ran forever. Those are the failures that quietly wreck cost and reliability in production.
How it works
You capture the full run as a sequence of steps, then score it against what a good path looks like. Some checks are exact: did it call the right tools with the right arguments, did it stay under a step budget, did it avoid repeating itself. Others are judged, often by an LLM rated against a reference trajectory: was each step a sensible move given what the agent knew at that point. This sits alongside task completion rate and tool-call accuracy, which each measure one slice of what the trajectory shows in full.
A support agent is asked to refund an order. It gets the refund right, but the trace shows it looked up the customer, forgot the result, looked them up again, then called the refund tool twice and got saved by an idempotency key. Final-answer grading says pass. Trajectory evaluation flags the duplicate lookups and the double refund call as a real defect waiting to bite on the next order.