Toxicity scoring
Toxicity scoring rates a piece of text for how hostile or harmful it is: slurs, threats, harassment, or abuse. The score is usually a number from 0 to 1, and you set a threshold above which the text gets blocked or flagged. The hard part is context. The same words can be an attack in one message and a quote or a reclaimed term in another.
Why it matters
A model that helps a user write a cruel message, or a support bot that snaps back at an angry customer, is a headline waiting to happen. Toxic output erodes trust instantly and can carry real legal and brand cost. It also runs the other way. A scorer that is too aggressive will censor a doctor discussing symptoms or a user quoting the abuse they received. You want to catch real harm without punishing people for talking about hard things.
How it works
A separate classifier reads the text and returns a toxicity probability, often broken into categories like insult, threat, and sexual content. Common tools are Google's Perspective API and the moderation endpoints shipped by model providers. You pick a threshold from your own labeled examples, then measure it like any classifier: track false positives (safe text flagged) and false negatives (toxic text missed) on a held-out set. Run it on both user input and model output, because a clean prompt can still trigger an ugly reply.
A gaming community's support bot starts mirroring the tone of the players it talks to, and one reply calls a frustrated user an idiot. A toxicity scorer on the output catches the reply at 0.92, blocks it, and swaps in a neutral apology before it ever reaches the screen. The rude draft lands in a review queue instead of the chat.