EN·ES·PT
AI agent testing · how to test LangGraph & agentic systems

AI agent testing: what breaks in production.

Agents don't fail the way functions fail. A unit test asserts one input against one output; an agent plans, calls tools, remembers, and loops — so the bugs live between turns, in the tool boundary, and in the parts that aren't deterministic. This page maps the failure surface that unit tests can't see, and the probes that catch each one before your users do.

The lesson · why single-turn evals aren't enough

A green single-turn suite hid a 502 on every follow-up.

On one build, a scoping chatbot passed its single-turn checks cleanly: ask a question, get a good answer, assertion green. But the moment a user sent a second message in the same conversation, the endpoint returned a 502 — every time. Single-turn testing can't see that, because it never sends turn two. A multi-turn golden probe — a scripted conversation that replays several turns and asserts on the whole trajectory — caught it immediately. That's the core of agent testing: you test the conversation and the loop, not one call.

Real lesson from a shipped build · no metric claimed beyond the failure itself
The failure surface

What unit tests miss — and what catches it.

Each row is a failure mode that a single-turn unit test structurally cannot observe, paired with the probe that does.

Agent failure surface → how a unit test misses it → the probe that catches it
Failure mode How a unit test misses it The probe that catches it
Multi-turn drift Asserts one input/output pair; the conversation never reaches turn two, so state that only corrupts across turns is invisible. Multi-turn golden probe — a scripted N-turn conversation asserting on the full trajectory, not one reply.
Tool-call errors Mocks the tool and checks the happy path; never exercises malformed args, wrong tool selection, or a tool that returns an error. Tool-boundary assertions — verify the agent picks the right tool, with valid args, and recovers when the tool fails.
Prompt injection across turns Tests trusted inputs only; hostile content planted in turn one that hijacks turn three is out of scope. Red-team battery — injected payloads in tool output and history, asserting the agent refuses and stays in policy.
Non-determinism & flakiness A single pass looks green; the same input silently produces a different, wrong path on the next run. Repeat / robustness runs — the same probe run K times, scored on pass rate, with paraphrase and retry checks.
Loop & termination faults No loop is executed, so runaway tool loops, early exits, and stuck plans never surface. Trajectory bounds — assert step count, termination, and that the agent reaches a valid stop state.
Ungrounded / uncited output Checks the text reads well, not whether every claim traces to a retrieved source. Citation-coverage gate — every generated claim must map to a source, enforced by construction.

Method, not vendor lock-in. The same three families — multi-turn golden probes, a red-team battery, and repeat/robustness runs — apply to LangGraph, a custom loop, or an off-the-shelf agent framework.

The loop, instrumented

Where the probes attach.

An agent is a loop: take input, plan, call a tool, respond — and often repeat. Every edge of that loop is a place a probe attaches. The gate at the end decides ship or block.

Agent loop with probe attachment points Input flows into Plan, then Tool call, then Respond, looping back to Plan for the next turn. Multi-turn golden probes attach across turns, tool-boundary assertions attach at the tool call, and a red-team battery injects at input and tool output. The trajectory is scored against an eval gate that either ships or blocks the release. Input user turn Plan choose next step Tool call act on the world Respond reply / continue next turn — where multi-turn drift hides tool-boundary assertions red-team inject scored trajectory Eval gate pass rate + red-team ship gates pass block — fails → back to fix
The loop is the unit under test. Probes attach at every edge; the eval gate turns a scored trajectory into a ship-or-block decision — the same discipline that let my QA OS block its own release when it found 15 high/critical CVEs.
The method · three families of probe

Multi-turn golden probes + a red-team battery + robustness.

1 · Multi-turn golden probes

Scripted conversations that replay several turns and assert on the whole trajectory — tool choice, state, and the final answer. This is the layer that would have caught the second-turn 502 on day one.

2 · Red-team battery
  • Injection planted in tool output and history, not just the prompt
  • Assert the agent refuses and stays in policy
  • Jailbreak, exfiltration, and scope-escape cases run every build
3 · Retry / robustness
  • Run each probe K times, score on pass rate — flaky ≠ passing
  • Paraphrase the same intent; the agent must hold
  • Bound loop steps and assert clean termination
The proof · not claims, receipts

This isn't theory — it's how I already work.

A gate that blocked its own release

My QA OS caught 15 high/critical CVEs, blocked the release, and was patched to 3,759 tests / 13-of-13 gates the same day. A gate that can't say "no" isn't a gate.

Grounded by construction

A RAG research dashboard with 100% citation coverage — no path exists to generate an uncited claim. The citation gate is structural, not a hope.

Deterministic suites that hold

A public SDET regression suite: 37/37 specs, 0 flakes, 15.3s. On live fintech workflows I took flake rate from 10% to under 1% across 500+ tests.

Public figures link to a receipt; the HighStrike numbers are from a prior role, self-reported · this site runs its own QA (100+ checks, axe-clean)
Related · more on testing AI

Worried about a specific agent shipping broken?

Book a call and we'll scope a probe suite for your agent — or paste your AI's answer into the live eval demo and watch it get graded right now.

Book a call → ▸ try the live eval demo LLM evaluation & QA →
frequently asked

Questions, answered.

Why are agents harder to test than a single prompt?
They take multi-step actions with tools and memory, so failures compound across steps and the same input can take different paths. You test the whole trajectory, not just the final answer.
What do you actually check on an agent?
Tool-call correctness, loop and termination behavior, guardrail adherence, cost and latency, and recovery from a bad step — plus the end-to-end outcome.
Does this work with LangGraph or a custom loop?
Yes. The same families of checks apply whether it’s LangGraph, a custom agent loop, or another framework; I wire the harness to your implementation.
Can you gate deploys on agent evals?
Yes. The agent eval runs in CI and blocks a release that regresses on the trajectories that matter.
Which agent frameworks do you support — LangGraph, CrewAI, or a custom loop?
All of them. The checks target the agent's trajectory, not the framework, so LangGraph, CrewAI, and hand-rolled loops are wired the same way. I adapt the harness to whatever runtime you already run.
How do you gate a deploy on agent behavior without blocking every release?
The eval runs in CI against a fixed set of golden trajectories and only blocks when behavior regresses on those. Non-determinism is handled with tolerance bands and repeated runs, so a healthy release ships and a real regression stops it.
What does a typical engagement timeline look like?
It starts with a short scoping pass to map your agent's critical trajectories, then a build phase for the golden probes and red-team battery, then handoff with the harness wired into CI. The work is front-loaded; the harness keeps running after I leave.
What do I actually receive at the end of the engagement?
A running test harness in your repo — golden probes, a red-team battery, and CI gating — plus a written report of the failure modes found and how each one is now caught. Everything is yours to run and extend without me.
How do you test agents that call external tools and APIs?
Tool calls are exercised against controlled fakes and recorded fixtures so tests stay deterministic, then spot-checked against the live integration. That catches malformed calls, bad arguments, and mishandled tool errors before they reach production.
How is an agent-testing engagement scoped and priced?
Scope follows your agent's surface area — the number of trajectories, tools, and guardrails that need coverage — and pricing is a fixed quote per agreed scope. You get the estimate before any work starts, with no open-ended hourly meter.