JTjason.teixeira() Glossary
Home / Learn / Glossary / Toxicity scoring
Safety

Toxicity scoring

Toxicity scoring rates a piece of text for how hostile or harmful it is: slurs, threats, harassment, or abuse. The score is usually a number from 0 to 1, and you set a threshold above which the text gets blocked or flagged. The hard part is context. The same words can be an attack in one message and a quote or a reclaimed term in another.

Why it matters

A model that helps a user write a cruel message, or a support bot that snaps back at an angry customer, is a headline waiting to happen. Toxic output erodes trust instantly and can carry real legal and brand cost. It also runs the other way. A scorer that is too aggressive will censor a doctor discussing symptoms or a user quoting the abuse they received. You want to catch real harm without punishing people for talking about hard things.

How it works

A separate classifier reads the text and returns a toxicity probability, often broken into categories like insult, threat, and sexual content. Common tools are Google's Perspective API and the moderation endpoints shipped by model providers. You pick a threshold from your own labeled examples, then measure it like any classifier: track false positives (safe text flagged) and false negatives (toxic text missed) on a held-out set. Run it on both user input and model output, because a clean prompt can still trigger an ugly reply.

In practice

A gaming community's support bot starts mirroring the tone of the players it talks to, and one reply calls a frustrated user an idiot. A toxicity scorer on the output catches the reply at 0.92, blocks it, and swaps in a neutral apology before it ever reaches the screen. The rude draft lands in a review queue instead of the chat.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy