frequently asked
Questions, answered.
What is an LLM evaluation, concretely?
A repeatable way to score your model’s output — accuracy, faithfulness, safety — against known-good cases, so a prompt or model change that makes things worse is caught before it ships.
How many test cases do you need to start?
Usually 50–200 representative cases — enough signal to gate on — expanded as real failure modes surface in production.
Can you gate our CI on eval results?
Yes. The eval runs in CI and blocks the merge if a threshold regresses (citation coverage, hallucination rate, injection resistance). Red before production, not after.
Which frameworks and models do you use?
Provider-agnostic — Promptfoo, DeepEval, and Pytest against a standard chat-completions interface, using your models and your accounts wherever possible.
How large does the golden dataset need to be to be useful?
A few dozen well-chosen cases beat hundreds of random ones. Start with the failure modes that actually hurt you — real user prompts, edge cases, past incidents — and grow the set as new failures surface.
Is LLM-as-judge scoring reliable enough to gate on?
Not on its own. Judge prompts are validated against human-labeled cases, and the gate leans on deterministic checks — exact match, citation presence, schema, refusal detection — wherever the answer allows, reserving the judge for the genuinely subjective calls.
How does the eval gate fit into the CI I already run?
It runs as one more check in your existing pipeline — a step in GitHub Actions or whatever you use — that passes or fails the build. No new platform to adopt; the eval lives beside your unit tests and blocks the merge on the same red X.
What happens to the gate when I swap models or edit a prompt?
That is exactly what it is for. Any model swap or prompt edit re-runs the full suite, and a change that regresses grounding, safety, or answer quality turns the build red before it reaches users.
Do you build this in my repo or a separate tool?
In your repo. The suite, the golden set, and the CI config are committed alongside your code, so you own it, run it, and extend it without depending on me or a hosted dashboard.
How long until the first quality gate is live?
The first working gate — a small golden set wired into CI and blocking on a real threshold — comes early, then the suite deepens from there. You see a build go red on a bad change before the engagement is over, not a slide about one.