EN·ES·PT
Hire · AI quality engineering

Hire an AI QA engineer before your LLM ships un-tested.

An AI feature that ships without tests doesn't crash — it fails quietly. It hallucinates a policy, cites a source that doesn't exist, leaks a prompt, or confidently gives one customer the wrong number. Nobody gets a stack trace. You find out from a screenshot on social media. I'm the engineer you bring in before that happens: I build the eval suites, judge gates, and CI regression discipline that catch a bad answer while it's still in your pipeline — not in front of your customers.

ISTQB CT-AI ISTQB Test Automation Engineer Live eval demo: grade your own AI ↗
Watch the full pitch
The risk · why this role exists

The people who build the feature are the wrong people to prove it works.

Whoever wrote the prompt already believes it works — that's the exact bias QA exists to break. An LLM feature has failure modes ordinary software doesn't: it's non-deterministic, it degrades silently when a model or prompt changes upstream, and "looks right" is not the same as "is right." Testing it takes a discipline that sits between building AI and traditional QA. That's the gap I fill.

What ships un-tested
  • Hallucinated facts, policies, and prices stated with total confidence
  • Citations to sources that do not exist
  • Prompt-injection and jailbreak paths nobody probed
  • Silent regressions when the model or prompt changes
What it costs you
  • A wrong answer to a real customer, in public
  • No way to tell if last week's fix broke this week's answer
  • "It worked in the demo" — and nowhere else
  • Trust you don't get a second chance to earn
What good looks like
  • A golden set of graded expected answers
  • Automated judge gates for safety and quality
  • A CI gate that blocks the release when a check fails
  • Every claim traceable to a source, by construction
The capability gap · who does what

An AI builder, a QA engineer, and me are not the same hire.

Most teams already have someone who can build an LLM feature, and maybe someone who can write UI tests. Neither one, on their own, closes the loop on AI quality. Here's the honest split.

Capability comparison
Capability AI builder Generalist QA Me (AI-QA)
Ships production LLM features Yes No Yes
Builds eval suites & judge gates Rarely Partial Yes
CI regression discipline Some Yes Yes
Safety / injection probing Rarely Rarely Yes
Certified in AI testing (ISTQB CT-AI) No Varies Yes

Yes = core strength · Partial / Varies = depends on the person · No / Rarely = not the role. Rated honestly — an AI builder who also builds eval gates exists, but it isn't the default.

The mechanism · the eval gate

Every AI answer runs a gauntlet before a customer sees it.

This is the shape of the system I build. Your AI's output isn't shipped on faith — it's scored against safety and quality checks, and a CI gate makes a binary call: ship it, or block the release and tell you exactly which check failed.

The AI eval gate flow AI output flows into safety and quality checks, then into a CI gate that either ships the release or blocks it. AI output the raw answer SAFETY + QUALITY • citation coverage • factual grounding • injection / jailbreak • format & schema • golden-set match CI gate pass? pass SHIP ✓ release proceeds fail BLOCKED ✗ release held

The same gate discipline runs this very site. Want to see a check score a real answer? Paste your AI's output into the live demo and watch it get graded.

▸ grade your own AI, live real eval · runs in your browser · no signup
Buyer's guide · how to tell a real hire from a resold-Zapier shop

Not everyone selling "AI QA" is testing AI.

A lot of "AI QA" offers are a Zapier flow with a chatbot bolted on. Here's how to tell the difference before you sign anything.

⚑ Red flag

Demos a chatbot answering questions and calls that "tested."

✓ Green flag

Shows you a failing eval and the exact check that caught it.

⚑ Red flag

Quotes an accuracy number with no golden set behind it.

✓ Green flag

Hands you a versioned golden set you own and can re-run.

⚑ Red flag

"It works" — but nothing runs in CI and nothing blocks a bad release.

✓ Green flag

Wires a gate into your pipeline that fails the build on regressions.

⚑ Red flag

Can't explain non-determinism, drift, or prompt-injection risk.

✓ Green flag

Probes safety and injection paths, and can prove the results.

The proof · claims I can back

The whole thesis of this site: proof over claims.

Every number here is linked to its evidence. If I can't show it, I don't say it.

My own QA OS blocked its own release

A QA system I built caught 15 high/critical CVEs, refused to ship itself, and was patched the same day to 3,759 tests passing across 13 of 13 gates. The before and after are both captured.

Regression discipline, measured

An SDET regression suite running 37 of 37 specs, 0 flakes, in 15.3s. The full source is public.

Grounded generation, by construction

A RAG research dashboard with 100% citation coverage by construction — there is no code path that can generate an uncited claim.

Track record on real systems
  • HighStrike (fintech): flake rate 10% → under 1%, 500+ tests on live trading workflows
  • The Home Depot (Fortune 50): Selenium framework for 2,300+ stores, regression 4h → 75min
  • This site runs its own QA — 100+ checks, axe-clean
public figures link to a receipt; the HighStrike & Home Depot numbers are from prior full-time roles, self-reported · nothing rounded up for a pitch
Related · more on testing AI

Worried about an AI feature you're about to ship?

Book a call and we'll scope the eval gate your feature needs — or grade one of your AI's answers live, right now, and see the discipline for yourself.

Scoped to your budget, not headcount — from a quick fix to a full build. Get a custom quote in 2 min →

Book a call → Grade your own AI live

Prefer to read first? See LLM evaluation & QA or the eval-gate guide.

frequently asked

Questions, answered.

Do I need a full-time hire, or can this be a contract?
Most engagements start as a scoped contract — an audit or a build — not a headcount. You get the coverage without hiring, and everything lands in your repo so your team can run it after handoff.
What does an AI QA engineer actually deliver?
The tests and evals that prove your AI feature works: golden sets, LLM-as-judge scoring, and CI gates that block a bad release, plus a runbook your team can operate.
How is this different from a normal QA/SDET?
Standard QA checks deterministic software. AI features are non-deterministic, so they need evaluation — is the answer faithful, cited, safe? — on top of traditional testing. I do both halves.
How fast can you start and see results?
A one-week audit gives you a prioritized plan and the highest-leverage gaps; a first working eval or CI gate typically lands within about two weeks.
Is this a contract engagement or a full-time hire?
A contract, by default — scoped to an outcome, not a seat. If you later want ongoing coverage we can move to a retainer, but nothing forces a headcount decision.
How quickly can you ramp on our stack and codebase?
Ramp starts on day one: I read your repo, prompts, and existing tests before writing anything. Fitting the way your stack already works matters more than imposing a new one.
Which tools and frameworks do you work in?
Pytest and Playwright for the test layer, LangGraph, CrewAI, and n8n for agent and workflow orchestration, and LLM-as-judge harnesses for the eval layer. I work against your model rather than swapping in my own.
Do you work directly in our repo and CI?
Yes — everything lands as pull requests in your repo and gates that run in your CI, not a side system you can't see. When the engagement ends, your team owns and runs all of it.
How do you collaborate remotely and across time zones?
Fully remote and async-first: work moves through pull requests, written updates, and a shared runbook so nothing depends on a live call. I keep overlap hours for review and unblock decisions.
How is the engagement scoped and priced?
Scope is fixed to a deliverable such as an audit, an eval suite, or a CI gate, and priced per engagement rather than by the hour. You approve the scope before work starts, so there's no open-ended meter.