EN·ES·PT
free artifact — no email wall

The LLM Feature Pre-Launch Eval Checklist

Eighteen checks before your LLM feature meets a customer — distilled from an 85-runner QA platform with ten dedicated AI-safety evals. If you can tick every box, ship. If you can't, you know exactly what to build first.

Ground truth

without this, every other check is opinion

Measurement

"it seems better" is not a metric

Safety battery

the failure modes that end up in screenshots

The gate

a red eval that blocks nothing is decoration

Blast-radius control

autonomy is earned with evidence, never assumed

Can't tick every box?

That's the gap I close in a 4-week fixed-scope engagement — the suite, the gate, and the runbook, in your repo.

How I build this → Download the PDF