The failure surface
What unit tests miss — and what catches it.
Each row is a failure mode that a single-turn unit test structurally cannot observe, paired with the probe that does.
Method, not vendor lock-in. The same three families — multi-turn golden probes, a red-team battery, and repeat/robustness runs — apply to LangGraph, a custom loop, or an off-the-shelf agent framework.
frequently asked
Questions, answered.
Why are agents harder to test than a single prompt?
They take multi-step actions with tools and memory, so failures compound across steps and the same input can take different paths. You test the whole trajectory, not just the final answer.
What do you actually check on an agent?
Tool-call correctness, loop and termination behavior, guardrail adherence, cost and latency, and recovery from a bad step — plus the end-to-end outcome.
Does this work with LangGraph or a custom loop?
Yes. The same families of checks apply whether it’s LangGraph, a custom agent loop, or another framework; I wire the harness to your implementation.
Can you gate deploys on agent evals?
Yes. The agent eval runs in CI and blocks a release that regresses on the trajectories that matter.
Which agent frameworks do you support — LangGraph, CrewAI, or a custom loop?
All of them. The checks target the agent's trajectory, not the framework, so LangGraph, CrewAI, and hand-rolled loops are wired the same way. I adapt the harness to whatever runtime you already run.
How do you gate a deploy on agent behavior without blocking every release?
The eval runs in CI against a fixed set of golden trajectories and only blocks when behavior regresses on those. Non-determinism is handled with tolerance bands and repeated runs, so a healthy release ships and a real regression stops it.
What does a typical engagement timeline look like?
It starts with a short scoping pass to map your agent's critical trajectories, then a build phase for the golden probes and red-team battery, then handoff with the harness wired into CI. The work is front-loaded; the harness keeps running after I leave.
What do I actually receive at the end of the engagement?
A running test harness in your repo — golden probes, a red-team battery, and CI gating — plus a written report of the failure modes found and how each one is now caught. Everything is yours to run and extend without me.
How do you test agents that call external tools and APIs?
Tool calls are exercised against controlled fakes and recorded fixtures so tests stay deterministic, then spot-checked against the live integration. That catches malformed calls, bad arguments, and mishandled tool errors before they reach production.
How is an agent-testing engagement scoped and priced?
Scope follows your agent's surface area — the number of trajectories, tools, and guardrails that need coverage — and pricing is a fixed quote per agreed scope. You get the estimate before any work starts, with no open-ended hourly meter.