JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Test Multi-Turn Agent Conversations
How-to guides

Test Multi-Turn Agent Conversations

Test the whole conversation, not just one reply, so state and memory bugs surface before users hit them.

A single-turn test checks one reply in isolation. But agents carry state across turns, so the bugs that hurt are the ones where turn 4 forgets what you said in turn 1. This guide builds a small Python test that drives a whole conversation and asserts on memory and state as it goes. Budget about thirty minutes.

#Before you start

  • Python 3.9 or newer with pytest and the openai package installed.
  • An agent you can call turn by turn (the example uses the OpenAI chat API).
  • An API key in an env var, for example OPENAI_API_KEY.

#Hold the conversation in one list

State lives in the message history. Keep a single list of messages and append every user turn and every reply to it, then pass the whole list back on the next call. This is the thing single-turn tests skip, and it is where memory bugs hide. Wrap it in a small helper so each test reads like a script.

conftest.py
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

class Conversation:
    def __init__(self, system):
        self.messages = [{"role": "system", "content": system}]

    def say(self, text):
        self.messages.append({"role": "user", "content": text})
        resp = client.chat.completions.create(
            model="gpt-4o-mini",
            messages=self.messages,
        )
        reply = resp.choices[0].message.content
        self.messages.append({"role": "assistant", "content": reply})
        return reply

#Write a test that spans turns

Give the agent a fact early, do some unrelated turns, then ask about the fact later. Assert on the later reply. The first reply does not matter here. If the agent held the state, it passes. If it dropped it, you catch the exact turn where memory failed.

test_memory.py
from conftest import Conversation

def test_remembers_name_across_turns():
    c = Conversation("You are a helpful assistant.")
    c.say("Hi, my name is Dana.")
    c.say("What is 2 plus 2?")            # distractor turn
    reply = c.say("What is my name?")
    assert "dana" in reply.lower()

#Test state that should change

Memory is one half. The other half is state that is supposed to update. Set a value, change it, then confirm the agent uses the new value and drops the stale one. This catches the agent that latches onto the first answer and never lets go.

test_state.py
from conftest import Conversation

def test_uses_updated_preference():
    c = Conversation("You are a travel assistant.")
    c.say("I want a window seat.")
    c.say("Actually, change that to an aisle seat.")
    reply = c.say("Which seat did I ask for?")
    text = reply.lower()
    assert "aisle" in text
    assert "window" not in text

#Grade the fuzzy turns with a judge

Keyword checks work for names and seats. For open replies, use a second model call as a judge with a tight rubric. Keep the rubric specific so a correct answer in different words still passes. Return a plain yes or no so the assert stays simple.

judge.py
from conftest import client

def judge(reply, rubric):
    r = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{
            "role": "user",
            "content": (
                f"Reply:\n{reply}\n\n"
                f"Does it satisfy this rule? {rubric}\n"
                "Answer only YES or NO."
            ),
        }],
    )
    return r.choices[0].message.content.strip().upper().startswith("YES")

#Run it

Run pytest from the terminal. Each test drives a full conversation and fails on the turn where state or memory broke. Once green, wire it into the same CI job as the rest of your evals so a regression stops at the pull request.

terminal
export OPENAI_API_KEY=sk-...
pytest -v

#Watch out for

  • Multi-turn tests are non-deterministic and cost more, because each test makes several model calls. Keep the suite small and focused on the state paths that actually matter, and run the big ones nightly instead of on every commit.
  • A long conversation can fail for a boring reason: the early turns fell out of the context window. That is a real bug your users will hit too, but do not confuse it with a reasoning failure. Check the token length before you blame the model.
  • If a test passes sometimes and fails sometimes, do not paper over it with a retry. A flaky memory test usually means the behavior is genuinely unreliable, which is exactly what you wanted to find.

#What you built

You have a test harness that carries a real conversation across turns and asserts on both remembered facts and updated state. It surfaces the memory and staleness bugs that single-turn tests cannot see. Grow it from real transcripts, especially the conversations where your agent forgot something and a user noticed.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms