free artifact — no email wall
The LLM Feature Pre-Launch Eval Checklist
Eighteen checks before your LLM feature meets a customer — distilled from an 85-runner QA platform with ten dedicated AI-safety evals. If you can tick every box, ship. If you can't, you know exactly what to build first.
Ground truth
without this, every other check is opinion
Measurement
"it seems better" is not a metric
Safety battery
the failure modes that end up in screenshots
The gate
a red eval that blocks nothing is decoration
Blast-radius control
autonomy is earned with evidence, never assumed
Can't tick every box?
That's the gap I close in a 4-week fixed-scope engagement — the suite, the gate, and the runbook, in your repo.