JTjason.teixeira() Glossary
Home / Learn / Glossary / Multi-turn evaluation
Agents

Multi-turn evaluation

Multi-turn evaluation tests a whole conversation across many back-and-forth turns, not just a single reply. Real chats have memory, follow-ups, and changing minds, so a bot that nails the first answer can still fall apart by turn five. A per-reply test scores each message on its own and misses the failure that only shows up once context piles up.

Why it matters

Most bugs in a real assistant live in the second half of the chat. The model forgets a constraint you gave three turns ago, contradicts itself, or loses the thread when a user changes their mind. Grade only single replies and you ship something that demos great, then frustrates anyone who actually talks with it. Users talk in conversations, so you have to test in conversations.

How it works

You script or simulate a full dialogue, sometimes with a second model playing the user, and grade the whole run. The checks look for what a single turn cannot show: did it remember the earlier constraint, stay consistent, recover after a correction, and reach the goal by the end. This overlaps with task completion rate for the outcome and agent trajectory evaluation for the path. The focus here is holding a coherent thread over time.

In practice

A user tells a support bot "I'm in Canada" on turn one, then four turns later asks about shipping cost. A single-turn test on that last question looks fine. A multi-turn test catches the bot quoting US rates because it dropped the Canada detail somewhere in the middle.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy