Prompt regression testing
Prompt regression testing means re-running your old test cases every time you change a prompt, so a fix in one place does not quietly break something else. You already had cases that passed. Now you check that they still pass. The trap is assuming a small wording tweak is safe. Prompts are sensitive enough that adding one sentence can flip answers you never touched.
Why it matters
Prompts are fragile in ways code is not. You tighten the instructions to fix one bad answer, ship it, and three other behaviors shift without warning. Without regression testing you hear about it from a user, or you never hear about it and quality slowly rots. Running the old cases turns "I think this is still fine" into "these 47 cases still pass and this one changed."
How it works
Keep a saved set of inputs with their known-good outputs, the same idea as a golden set. After any prompt edit, run the whole set and compare new outputs to the reference. For fixed answers, use exact or near-exact matching. For open-ended ones, lean on semantic similarity or an LLM judge. The signal you watch is the diff: which cases newly failed, and whether the case you were trying to fix actually got fixed.
A support bot gives a clumsy answer about shipping times, so you rewrite the prompt to be more concise. You re-run your saved cases and two refund questions that passed yesterday now give the wrong window, because "be concise" made the model drop a key detail. You caught it before it shipped, instead of after a customer acted on the bad answer.