Eval-driven development
Eval-driven development means you write the test for an AI feature before you build the feature. You define what a good answer looks like, encode it as an eval, and only then touch prompts or code. The catch: done stops meaning "it ran without errors" and starts meaning "it passed the eval." That is a much higher bar than most people are used to.
Why it matters
Without an eval up front, you tune your prompt until the three cases you happened to try look good, then ship and hope. You have no honest way to tell whether your next change helped or quietly broke something. Writing the eval first forces you to decide what success is before you fall in love with an implementation. It turns "looks good to me" into a number you can defend.
How it works
Start from the spec. Write a small set of inputs with expected behavior, pick how each one gets graded (exact match, semantic similarity, a rubric, or an LLM-as-a-judge for open-ended answers), and wire it into a harness that runs on every change. The feature is finished when the score crosses the bar you set. Every regression later becomes a new case in the set. Cheap versions run in a script. Serious ones run in CI so a bad change fails before merge.
You are building a refund bot. Before writing a single prompt, you add a case: "customer asks about a 40-day-old order, expected answer cites the 30-day policy and offers store credit." You run it, it fails, and now you have a target. You tweak the prompt until that case and thirty others pass, and then the feature counts as built.