Build a Tool-Call Accuracy Test Harness for AI Agents
Check that your agent reaches for the correct tool with the correct arguments, every time.
A tool-call accuracy harness runs your agent against known prompts and checks that it reaches for the right tool with the right arguments. Once it exists, a prompt tweak or model swap that quietly breaks routing fails a test instead of a user. This guide builds one with pytest, using a real model call and parametrized cases. Budget about forty minutes.
#Before you start
- Python 3.10+ and pytest installed (
pip install pytest openai). - An agent that exposes its tool schemas. The examples use the OpenAI tool-calling format.
- An API key in an env var (
OPENAI_API_KEY).
#Write down the tools and the expected routing
Before any test, list the cases as data: a prompt, the tool you expect, and the arguments that matter. Keep the expected args to the fields you actually care about, so a harmless extra field does not fail the case. Start with one clean case per tool plus a couple of tricky ones that sit near a boundary between two tools.
CASES = [
{
"id": "weather-basic",
"prompt": "What's the weather in Tokyo right now?",
"tool": "get_weather",
"args": {"city": "Tokyo"},
},
{
"id": "send-email",
"prompt": "Email jane@acme.com and say the report is ready.",
"tool": "send_email",
"args": {"to": "jane@acme.com"},
},
{
"id": "weather-not-email",
"prompt": "Should I bring an umbrella to Denver tomorrow?",
"tool": "get_weather",
"args": {"city": "Denver"},
},
]#Make one function that returns the chosen tool call
The harness needs a single call that takes a prompt and returns which tool the model picked and with what arguments. This wraps whatever your agent actually does. Here it is a plain model call with your tool schemas and tool_choice="auto", the same path your agent uses.
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
TOOLS = [
{"type": "function", "function": {
"name": "get_weather",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}}},
{"type": "function", "function": {
"name": "send_email",
"parameters": {"type": "object",
"properties": {"to": {"type": "string"},
"body": {"type": "string"}},
"required": ["to"]}}},
]
def pick_tool(prompt):
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
tools=TOOLS,
tool_choice="auto",
)
calls = resp.choices[0].message.tool_calls
if not calls:
return None, {}
call = calls[0]
return call.function.name, json.loads(call.function.arguments)#Turn each case into a pytest
Parametrize over the cases so each one shows up as its own named test. Assert the tool name exactly, and check that every expected argument is present and correct. Use a subset check on the args, so the model is free to add fields you did not pin down.
import pytest
from agent import pick_tool
from cases import CASES
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
def test_tool_call(case):
tool, args = pick_tool(case["prompt"])
assert tool == case["tool"], (
f"expected {case['tool']}, got {tool}"
)
for key, want in case["args"].items():
assert args.get(key) == want, (
f"arg {key}: expected {want!r}, got {args.get(key)!r}"
)#Run it and read the failures
Run pytest and let the case ids point you at exactly what broke. A wrong tool and a wrong argument are different failures, and the harness tells them apart. This is your loop while you tune the tool descriptions or the system prompt.
export OPENAI_API_KEY=sk-...
pytest -v
# test_tool_calls.py::test_tool_call[weather-basic] PASSED
# test_tool_calls.py::test_tool_call[send-email] PASSED
# test_tool_calls.py::test_tool_call[weather-not-email] FAILED#Gate it in CI
Add a workflow that runs the harness on every pull request. Pytest exits non-zero on any failure, and that fails the job. Put the API key in repo secrets, never in the file. Mark the job as a required check so a routing regression cannot merge by accident.
name: tool-call-accuracy
on: [pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install pytest openai
- run: pytest -v
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}#Watch out for
- Tool selection is not fully deterministic. A borderline prompt can flip between two reasonable tools, so keep those cases clearly one-sided or accept a small set of valid tools.
- Exact-matching every argument makes the suite brittle. Pin only the args that matter for correctness and let the rest vary, or you will chase phrasing instead of behavior.
- Every run costs a real API call. Keep the suite small and fast on pull requests, and run the larger set on a schedule if it grows.
#What you built
You now have a harness that checks your agent picks the right tool with the right arguments, running as named tests locally and as a required check in CI. It is small, it runs the real model path, and it catches routing breaks that are easy to miss by eye. Grow the case set from production, especially the prompts where the agent reached for the wrong tool.