EN·ES·PT
LLM evaluation consultant

The eval suite that decides whether your model ships.

An LLM-evaluation engagement isn't a report — it's a running system. I build the golden dataset, the scoring rubric, and the CI gate that grades every model or prompt change against it, then blocks the merge when quality regresses. You can watch a working version of that judge grade an answer live, right now, before you talk to me.

Live judge demo: eval.html → Grade your AI’s answer → ISTQB CT-AI + Test Automation Engineer
Watch · the eval wedge in ~45s
4
scoring dimensions per answer
2
signals: LLM-judge + deterministic checks
1
CI gate that can block a merge
100%
citation coverage on the RAG build*

*The RAG research dashboard cites 100% of its claims by construction — there is no code path that generates an uncited sentence. Verified on case-studies.html. Every other figure on this page is a design fact of the method, not a client average.

The rubric · what the live demo scores

Four dimensions, each with a catch and a score.

Every answer the judge sees is graded on the same four axes. Two are deterministic (they either pass or they don't); two use an LLM-as-judge with an explicit rubric so the score is explainable, not a vibe. This is the exact rubric behind the live demo.

Evaluation rubric — dimension × failure caught × scoring method
Dimension What it catches How it’s scored
Grounding Answers that drift from the source of truth — claims the provided context doesn’t actually support. LLM-as-judge against the supplied source, plus a deterministic citation check: every claim must trace to a passage.
Hallucination Confident invention — fabricated facts, fake figures, or citations to sources that were never given. Deterministic claim-vs-context diff surfaces unsupported statements; judge rates severity.
Safety & injection Prompt-injection compliance, leaked system instructions, and unsafe or out-of-policy output. Deterministic pattern & refusal checks on adversarial inputs; a fail here can hard-block ship.
Answer quality Technically-grounded but useless replies — off-intent, incomplete, or ignoring the actual question. LLM-as-judge on relevance, completeness and directness, scored against a written rubric.
▸ paste your AI’s answer & watch it get graded the demo runs this rubric on real input, live
The methodology · golden set → judge → gate

How a regression gets caught before it ships.

A model or prompt change enters CI. It runs against a curated golden set of representative cases. Each output is scored by the LLM judge and the deterministic checks. If any dimension falls below its floor, the gate blocks the merge — the change never reaches users. If everything clears, it ships.

LLM evaluation gate flow A change and its golden set feed an evaluation stage combining an LLM judge and deterministic checks, producing scores that hit a CI gate; passing changes ship, failing changes are blocked. Model / prompt change → CI Golden set curated cases EVALUATE LLM-as-judge + deterministic checks GATE score ≥ floor? SHIP ✓ all floors met BLOCKED ✕ regression
The gate is the deliverable. Without it, evals are a dashboard nobody reads; with it, a bad change physically cannot merge.
The proof · this gate has already blocked a release

A red → green story I can show you the receipts for.

On my own QA platform, nexural-qa-os, the gate did exactly what a gate is supposed to do: it caught 15 high/critical CVEs and blocked its own release. Same day, the issues were patched and the system went green — 3,759 tests passing, 13 of 13 gates clear. The evidence is two captured runs, before and after.

Before — release blocked
  • 15 high/critical CVEs detected
  • Gate refused to pass the build
  • Nothing shipped
After — same day, green
  • 3,759 tests passing
  • 13 / 13 gates clear
  • Patched, re-run, shipped
Why it matters to you

A gate only earns trust the day it stops your own release. This one did — and I kept the captures instead of just claiming it.

▸ the blocked run ▸ the green run two captured runs · before & after · same day
Adjacent proof · the same discipline, elsewhere

Eval rigor is one facet of how I test.

Deterministic suites

A public SDET regression suite: 37/37 specs, 0 flakes, 15.3s. Green means green.

Flake under control (fintech)

HighStrike (a prior full-time role, self-reported): flake rate driven from 10% to under 1% across 500+ tests on live trading workflows.

Scale (Fortune 50)

The Home Depot: a Selenium framework for 2,300+ stores; regression cut from 4h to 75min.

▸ the full case studies and yes — this site runs its own QA: 100+ checks, axe-clean
Related · more on testing AI

Worried about one specific AI feature?

Grade it live in 30 seconds, or book a call and we’ll scope the golden set, the rubric, and the gate that protects it.

Book a call → ▶ try the live judge the full service →
frequently asked

Questions, answered.

What is an LLM evaluation, concretely?
A repeatable way to score your model’s output — accuracy, faithfulness, safety — against known-good cases, so a prompt or model change that makes things worse is caught before it ships.
How many test cases do you need to start?
Usually 50–200 representative cases — enough signal to gate on — expanded as real failure modes surface in production.
Can you gate our CI on eval results?
Yes. The eval runs in CI and blocks the merge if a threshold regresses (citation coverage, hallucination rate, injection resistance). Red before production, not after.
Which frameworks and models do you use?
Provider-agnostic — Promptfoo, DeepEval, and Pytest against a standard chat-completions interface, using your models and your accounts wherever possible.
How large does the golden dataset need to be to be useful?
A few dozen well-chosen cases beat hundreds of random ones. Start with the failure modes that actually hurt you — real user prompts, edge cases, past incidents — and grow the set as new failures surface.
Is LLM-as-judge scoring reliable enough to gate on?
Not on its own. Judge prompts are validated against human-labeled cases, and the gate leans on deterministic checks — exact match, citation presence, schema, refusal detection — wherever the answer allows, reserving the judge for the genuinely subjective calls.
How does the eval gate fit into the CI I already run?
It runs as one more check in your existing pipeline — a step in GitHub Actions or whatever you use — that passes or fails the build. No new platform to adopt; the eval lives beside your unit tests and blocks the merge on the same red X.
What happens to the gate when I swap models or edit a prompt?
That is exactly what it is for. Any model swap or prompt edit re-runs the full suite, and a change that regresses grounding, safety, or answer quality turns the build red before it reaches users.
Do you build this in my repo or a separate tool?
In your repo. The suite, the golden set, and the CI config are committed alongside your code, so you own it, run it, and extend it without depending on me or a hosted dashboard.
How long until the first quality gate is live?
The first working gate — a small golden set wired into CI and blocking on a real threshold — comes early, then the suite deepens from there. You see a build go red on a bad change before the engagement is over, not a slide about one.