JTjason.teixeira() Glossary
Home / Learn / Glossary / Perplexity
Evaluation

Perplexity

Perplexity is a number that says how surprised a language model was by a piece of text. Low perplexity means the model found the words predictable; high means the words caught it off guard. The catch: it only tells you how well the model predicts text, not whether that text is correct or useful.

Why it matters

It is one of the few eval numbers you can compute automatically, with no human and no judge model in the loop, so it is cheap to run on every checkpoint. That makes it great for catching regressions during training or fine-tuning: if perplexity on your held-out text suddenly jumps, something broke. It is a poor judge of a finished product, though, because a model can be fluent and confident while being confidently wrong.

How it works

You take text the model has never seen, feed it in, and ask the model what probability it assigned to each actual next word. Perplexity is the average of those probabilities, flipped so that lower is better. A perplexity of 10 roughly means the model was as unsure as if it were guessing between 10 equally likely words at each step. It only works on models that expose token probabilities, so it fits open-weight models better than a closed API.

In practice

You fine-tune a small model on your company's support transcripts. Before training, its perplexity on a held-out batch of real tickets is 45. After training it drops to 12, which tells you the model now finds your ticket style far less surprising. That number does not promise good answers, so you still run a golden set on top of it.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy