AI agent testing →
Testing agents that plan, call tools, and take multi-step actions.
Testing systems that take actions, not just answer questions. Making multi-step, tool-using agents trustworthy: trajectory evaluation, tool-call accuracy, multi-turn testing, and the non-determinism problem.
Testing agents that plan, call tools, and take multi-step actions.
A multi-agent audit gauntlet — and the false error one agent was teaching.
Coordinating parallel agents against one design file without merge chaos.
Build a harness that tests an agent picks the right tool and arguments.
Test multi-turn agent conversations, including memory and state.
The production failure modes of AI agents and how to catch them early.
How to evaluate the whole path an agent takes, not just its final answer.
How to test non-deterministic AI without chasing noise.
How to set a real reliability bar for an AI agent, tied to its blast radius.