Non-determinism
Non-determinism is when the same prompt, sent to the same model, can come back different each time. Nothing broke. Language models pick each word from a set of probabilities, and that dice-roll means today's answer is not guaranteed to match yesterday's. It is the thing that makes AI features feel unlike normal code, where the same input always gives the same output.
Why it matters
If you test a feature once, see it pass, and ship it, you have proven it works that one time. The same question can pull a clean answer for you and a broken one for a customer an hour later. This is also why a bug can be real and still refuse to reproduce on demand, which makes debugging feel like chasing ghosts. You cannot trust a single green run.
How it works
You reduce it by lowering the sampling randomness. Set temperature toward 0 so the model favors its most likely token. Output gets steadier, though never perfectly identical. Handle the rest by testing for a range of acceptable answers instead of one fixed string, running each case several times, and scoring the pass rate. A golden set run many times tells you how often a case holds up.
A code generator is asked to write the same function three times. Once it returns clean, working code. Once it adds an import nobody requested. Once it renames a variable and quietly breaks a test. Same prompt, three outputs, and only running it many times reveals which behavior is the real one.