JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Build a Tool-Call Accuracy Test Harness for AI Agents
How-to guides

Build a Tool-Call Accuracy Test Harness for AI Agents

Check that your agent reaches for the correct tool with the correct arguments, every time.

A tool-call accuracy harness runs your agent against known prompts and checks that it reaches for the right tool with the right arguments. Once it exists, a prompt tweak or model swap that quietly breaks routing fails a test instead of a user. This guide builds one with pytest, using a real model call and parametrized cases. Budget about forty minutes.

#Before you start

  • Python 3.10+ and pytest installed (pip install pytest openai).
  • An agent that exposes its tool schemas. The examples use the OpenAI tool-calling format.
  • An API key in an env var (OPENAI_API_KEY).

#Write down the tools and the expected routing

Before any test, list the cases as data: a prompt, the tool you expect, and the arguments that matter. Keep the expected args to the fields you actually care about, so a harmless extra field does not fail the case. Start with one clean case per tool plus a couple of tricky ones that sit near a boundary between two tools.

cases.py
CASES = [
    {
        "id": "weather-basic",
        "prompt": "What's the weather in Tokyo right now?",
        "tool": "get_weather",
        "args": {"city": "Tokyo"},
    },
    {
        "id": "send-email",
        "prompt": "Email jane@acme.com and say the report is ready.",
        "tool": "send_email",
        "args": {"to": "jane@acme.com"},
    },
    {
        "id": "weather-not-email",
        "prompt": "Should I bring an umbrella to Denver tomorrow?",
        "tool": "get_weather",
        "args": {"city": "Denver"},
    },
]

#Make one function that returns the chosen tool call

The harness needs a single call that takes a prompt and returns which tool the model picked and with what arguments. This wraps whatever your agent actually does. Here it is a plain model call with your tool schemas and tool_choice="auto", the same path your agent uses.

agent.py
import json
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

TOOLS = [
    {"type": "function", "function": {
        "name": "get_weather",
        "parameters": {"type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]}}},
    {"type": "function", "function": {
        "name": "send_email",
        "parameters": {"type": "object",
            "properties": {"to": {"type": "string"},
                           "body": {"type": "string"}},
            "required": ["to"]}}},
]

def pick_tool(prompt):
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        tools=TOOLS,
        tool_choice="auto",
    )
    calls = resp.choices[0].message.tool_calls
    if not calls:
        return None, {}
    call = calls[0]
    return call.function.name, json.loads(call.function.arguments)

#Turn each case into a pytest

Parametrize over the cases so each one shows up as its own named test. Assert the tool name exactly, and check that every expected argument is present and correct. Use a subset check on the args, so the model is free to add fields you did not pin down.

test_tool_calls.py
import pytest
from agent import pick_tool
from cases import CASES

@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
def test_tool_call(case):
    tool, args = pick_tool(case["prompt"])

    assert tool == case["tool"], (
        f"expected {case['tool']}, got {tool}"
    )
    for key, want in case["args"].items():
        assert args.get(key) == want, (
            f"arg {key}: expected {want!r}, got {args.get(key)!r}"
        )

#Run it and read the failures

Run pytest and let the case ids point you at exactly what broke. A wrong tool and a wrong argument are different failures, and the harness tells them apart. This is your loop while you tune the tool descriptions or the system prompt.

terminal
export OPENAI_API_KEY=sk-...
pytest -v
# test_tool_calls.py::test_tool_call[weather-basic] PASSED
# test_tool_calls.py::test_tool_call[send-email] PASSED
# test_tool_calls.py::test_tool_call[weather-not-email] FAILED

#Gate it in CI

Add a workflow that runs the harness on every pull request. Pytest exits non-zero on any failure, and that fails the job. Put the API key in repo secrets, never in the file. Mark the job as a required check so a routing regression cannot merge by accident.

.github/workflows/tool-calls.yml
name: tool-call-accuracy
on: [pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install pytest openai
      - run: pytest -v
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

#Watch out for

  • Tool selection is not fully deterministic. A borderline prompt can flip between two reasonable tools, so keep those cases clearly one-sided or accept a small set of valid tools.
  • Exact-matching every argument makes the suite brittle. Pin only the args that matter for correctness and let the rest vary, or you will chase phrasing instead of behavior.
  • Every run costs a real API call. Keep the suite small and fast on pull requests, and run the larger set on a schedule if it grows.

#What you built

You now have a harness that checks your agent picks the right tool with the right arguments, running as named tests locally and as a required check in CI. It is small, it runs the real model path, and it catches routing breaks that are easy to miss by eye. Grow the case set from production, especially the prompts where the agent reached for the wrong tool.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms