Agent Trajectory Evaluation: Testing More Than the Answer
An agent can reach the right answer through a broken path. Trajectory evaluation grades the path, not just the destination.
An agent can land on the right answer after a completely broken run. It called the wrong tool, ignored the result, retried the same call three times, then guessed correctly. If you only check the final output, that run passes. Trajectory evaluation grades the path the agent took to get there, which is where the real failures and the real costs live.
#Why the final answer lies to you
Final-answer evaluation treats an agent like a function: input in, output out, compare to expected. That works for a model that answers in one shot. It falls apart for an agent, because an agent is a sequence of decisions, and a lucky sequence and a sound sequence produce the same last line.
Here is a concrete case. Ask a support agent to refund an order. It searches for the order, gets a permission error, silently retries the same call, fails again, falls back to a hardcoded default, and reports success. The customer got their refund, so the output matches. The trajectory shows an agent that cannot handle a permission error and papers over the failure with a guess. Ship enough of those and you learn about the broken path in production, one wrong refund at a time.
#What a trajectory actually contains
A trajectory is the ordered log of what the agent did between question and answer: each tool call, the arguments it passed, the result it got back, and the reasoning step that led to the next call. That is the thing you grade.
The useful checks are specific. Did it call the right tools? A weather question that never calls the weather API is suspect even if the answer is right. Did it call them in a sane order? Booking a flight before checking the date is a red flag. Did it react to what came back? An agent that gets an empty search result and keeps going as if it found something is narrating, not reasoning. Did it waste steps? Twelve tool calls for a two-call task is a cost and latency bug even when the answer is correct.
#How to grade the path without drowning
Two approaches, and you usually want both. The first is a reference trajectory. For a fixed test case, write down the tool sequence a competent run should produce, then compare the real run against it. This is exact and cheap to check. It is also brittle, because there is often more than one correct path, and a strict match will fail a run that was actually fine.
The second is a rubric applied by an LLM judge reading the whole trace. You hand the judge the goal and the full trajectory and ask targeted questions: was every tool call justified by a real need, did the agent recover from errors, were there redundant or contradictory steps. This handles multiple valid paths, but it inherits every judge weakness, so calibrate it against traces you graded by hand before you trust it. A practical split: reference matching for the handful of flows with one obvious correct path, rubric judging for the open-ended ones.
#What this does not catch
Trajectory evaluation tells you the path was sound. It does not tell you the answer was right. A perfectly reasoned run can call the correct tools in the correct order and still return a wrong number because a tool gave bad data. You need both graders. The trajectory catches process failures the output hides. The output catches correctness failures a clean-looking path hides.
It is also more work to maintain. Reference trajectories break when you change a tool signature or reorder a step, and you will spend real time deciding whether a failed match is a regression or just a new valid path. That maintenance cost is the honest tradeoff. It is worth paying for agents that take irreversible or expensive actions, and probably not worth it for a read-only agent where a wrong answer is cheap to catch and easy to undo.
#The bottom line
Trajectory evaluation does not replace checking the final answer. It is the other half. Keep output grading for correctness, add trajectory grading for the agents where how they got there actually matters, and accept that the reference traces cost something to keep current. For anything that spends money or takes an action you cannot easily reverse, that cost is the cheapest insurance you will buy.