AI automation × QA / LLM eval engineerN 28.5384° W 81.3789°
I ship AI features. Then I prove they work.
I'm who you bring in when you're about to ship an LLM feature that could embarrass you in front of a customer. I build it, and the gate that catches it the moment it drifts. see why →
LLM apps and agents, plus workflow automation (RAG, LangGraph, n8n/Make), paired with the regression and eval harness that keeps them honest in production: Playwright, Pytest, k6, Promptfoo, DeepEval, LLM-as-judge. Most people build the AI or test it. I do both, and every claim below links to an artifact.
13/13 proof gates · red→green today
3,759 tests passing · 91% line coverage
verified · reproducible by one command
one engagement at a time — currently booking Q4 2026
fig. 01: a feature fails eval → it goes back. that's the whole point.
the 30-second version
In a sentence: I build your AI feature — then I prove it works.
what I do
LLM & RAG evaluation, test automation & CI, and workflow automation — the AI feature and the evidence it holds up.
why me
13 years in software quality at The Home Depot and in fintech. Today I build and run two live AI products solo — and this site runs its own QA in public.
how to start
A one-week AI Quality Audit. You leave with a plan either way, and the fee credits into the build if we continue.
Your operation today, into a system you can prove.
Scroll it. The same five-stage flow, from manual-and-unprovable to automated-and-eval-gated. The shape of what I build for you.
built & operated — not client logos, my own products
Two live products. Built and run by myself.
The strongest proof I can offer isn't a testimonial — it's software I ship and operate solo, every day. Two products on one engineering system. Open the front door of either.
Agent warehouse monorepo: voice agents, outbound SDR, RAG support chat, and workflow orchestrators on Mastra + LangGraph — with a multi-provider LLM router, cost tracking per run, Langfuse tracing, and a hard TCPA consent gate on every outbound call.
blank page → staged client agent in days, not weeks — estimated from template scaffold time
MastraLangGraphVercel AI SDKSupabaseLangfuse
Private — walkthrough on request
~/sage-agents
$ pnpm tsx infra/scripts/create-agent.ts \
inbound-voice acme-inbound-voice
scaffolded @sage/agent-acme-inbound-voice
✓ prompts · tools · evals/golden.json ready
fig. 02 — architecture · live output
02
Autonomous workflowVERIFIED
sage-kernel →
Proof-first MCP engineering OS: 140 tools an AI agent drives through policy, signed approvals, and a hash-chained proof ledger. Nothing is "done" because a model said so — a claim-firewall rejects unproven success language.
140 MCP tools · 78 release gates — verified: gate suite + hash-chained proof ledger in the public repo
RAG knowledge base where every answer is extractive and cited. Durable ingestion worker, pgvector retrieval, persisted eval runs, and a full audit trail from source to answer.
100% citation coverage — verified 2026-08-15: ran it live, asked a real question, screenshot below is the actual answer
FastAPIpgvectorGeminiReact
Private — walkthrough on request
fig. 04 — architecture · real UI, real query, 2026-08-15 · click to zoom
04
QA framework + CIVERIFIED
playwright-sdet-regression-suite →
Release-critical e-commerce flows under regression with the evidence a release manager would ask for: traces, screenshots, four reporters, and a written risk model. CI uploads artifacts on every push.
37/37 specs · 15.3s · 0 flakes — verified: evidence/ folder in the public repo, run dated 2026-07-10
85 quality runners under one CLI — including hallucination, jailbreak, prompt-injection, toxicity, and PII-leak evals for LLM features. Every score is computed from a real command and packaged as ed25519-signed evidence.
3,759 tests · 91% coverage · 13/13 — verified 2026-08-15: the CVE gate went red at 11:39, blocked release, and was green by 14:35. Both runs published.
Answer 3 questions or talk to Nadine, my AI associate — you'll get an itemized plan with a price range in about 2 minutes. No signup, no sales call required.
✓ Prefer to start small? Begin with a fixed-price $497 audit week — credited straight into the build if you continue.
How the work actually went, as spec sheets you can audit.
Outcome first; the full eight-section spec — including the parts most case studies hide — is one click deeper. Every headline number states how it was measured.
spec BRIEF/01 · rev A
BRIEF/01
Feedback triage that runs itself
Make + Gemini workflow
Negative feedback now surfaces in seconds — not at the Friday review.
Customer feedback arrived through a form and sat unread. Negative signals surfaced days late, after the customer was already gone.
Measured results
Negative feedback surfaces in real time instead of at end-of-week review — measured as webhook-to-alert latency of seconds vs a manual weekly pass.
▸view full spec — 6 more sections
Why not a chatbot
Nobody wants to converse with their own feedback queue. The value is deterministic: classify, summarize, route, alert. A chat interface would add latency and remove auditability.
Architecture
Make scenario, four steps end to end — see fig. 1. Each step re-runnable in isolation.
Retrieval / memory
None needed — each message is classified independently. Deliberately avoided: statelessness keeps the pipeline debuggable.
Human approval points
The AI never replies to a customer. It routes to a human with a summary; the alert email is the handoff, not the resolution.
Production safeguards
Structured output schema on the LLM step, fallback label on parse failure, and the raw message always logged alongside the classification.
What I would harden next
A golden set of labeled messages run nightly through Promptfoo to catch classification drift when the model version changes.
stack — Make · Gemini 2.5 Flash · Google Forms · Sheets · Gmail
spec BRIEF/02 · rev A
BRIEF/02
RAG that cites or shuts up
ai-research-dashboard
100% of answers cite their source. The pipeline has no uncited path.
Research teams need answers from their own corpus — but a RAG system that paraphrases confidently without sources is a liability, not a tool.
Measured results
100% of answers citation-backed — verified by design: there is no uncited generation path. Quality, safety, latency, and cost tracked per query on the analytics endpoint.
▸view full spec — 6 more sections
Why not a chatbot
A chat wrapper over the corpus was the obvious build. The actual need was an operations dashboard: sources, ingestion jobs, eval runs, and feedback all inspectable — the Q&A is one feature inside it.
Architecture
FastAPI API + durable ingestion worker feeding fig. 1; provider gateway swaps local deterministic embeddings for Gemini per environment. React dashboard on top.
Retrieval / memory
Persisted queries, answers, and citations. Deleting a source cascades: chunks, embeddings, jobs, payloads, citations — no orphaned memory.
Human approval points
Answer-level feedback endpoint feeds a review loop; eval runs are triggered and persisted by API so a human signs off before provider or threshold changes ship.
Production safeguards
Extractive-only answers (cite or abstain), full audit trail on source/query/eval/feedback events, deterministic local providers for reproducible dev, Gemini fallback-to-local on missing keys.
What I would harden next
Promote eval thresholds into deployment gates so a retrieval regression blocks the release, not the retro.
Quality dashboards state numbers they can’t defend. For LLM features it’s worse: teams ship prompt changes with no regression signal at all.
Measured results
3,759 tests, 91.12% line coverage, 13/13 gates. On 2026-08-15 the CVE gate went honestly red (15 high/critical in deps) and blocked its own release — patched and PROVEN again the same afternoon, both runs published verbatim. One command regenerates the scorecard.
▸view full spec — 6 more sections
Why not a chatbot
An "AI QA assistant" that opines on quality is exactly the failure mode. The verdict must be computed by code from command exit codes — an LLM never grades its own homework here.
Architecture
One CLI driving fig. 1: every runner is a plugin (detect → plan → run → collectEvidence); one hung runner can’t stall a run. Two gates: <15ms pre-commit, full-repo merge gate.
Retrieval / memory
Evidence ledger instead of memory: every run’s exact command, exit code, and parsed metric recorded to a proof ledger, re-verifiable offline.
Human approval points
The autonomous fix loop is bounded. It fixes only what the harness can verify, commits on improvement, reverts on regression, stops on an honest stall for human review.
Production safeguards
Anti-hallucination contract: no score without a backing artifact. Ratcheting coverage floors, honest-skip discipline, ed25519-signed redacted evidence.
LLM-eval coverage
Ten dedicated AI-safety runners: bias, consistency, hallucination, jailbreak, prompt-injection, refusal, toxicity, PII-leak, cost, latency — the harness I bring to client LLM features.
The market filled up fast with shops wrapping a rented Zapier flow in a markup. So don't take my word for it — here's the checklist I'd hand you to vet anyone, me included. Run every line on me.
✗ what most "AI automation" shops do
✓ what you get here
"Book a call" and a vague promise — nothing to inspect before you commit.
A live, clickable demo and public receipts before you ever talk to me. Every number on this page links to the verbatim run.
A resold Zapier/Make flow that looks identical for every client.
Real code in your repo, fit to your stack. You own it at handoff — nothing locked to my platform.
Guaranteed ROI, quoted before they've even seen your system.
A fixed-price, one-week audit that tells you the truth first — and credits straight into the build if you continue.
NDA'd "case studies" and round numbers you can't check.
Receipts you can open: 37/37 specs green in CI, 10%→<1% flake, 15→0 CVEs — each one linked to the run.
Ships the AI feature and hopes it behaves in front of your customers.
An eval battery + CI gate — faithfulness, injection, toxicity, regression — so it's tested before it touches a customer, not after.
You're locked in; they become the single point of failure.
Runbook, walkthrough, documented handoff. Your team runs it without me — and the work survives me being unavailable.
watch · 45 seconds
See the eval battery block a release — before a customer would.
Ten checks, scored by code. One fails, the gate stops the ship. Then the real caught-blocked-proven arc from my own repo.
Fixed scope, fixed deliverables, evidence included. Each offer is a productized version of something in the briefs above.
Hiring full-time? These are the three capabilities I'd bring to your team on day one, each backed by a shipped system above.
LLM feature QA & eval harness
Your LLM feature gets a regression suite: golden traces, LLM-as-judge scoring for faithfulness and safety, and a CI gate that blocks the deploy when quality drops.
→ Promptfoo / DeepEval suite on your real traffic patterns
A risk-scoped Playwright/Pytest suite with the evidence discipline of a release manager: traces, reports, artifacts — wired into GitHub Actions from day one.
→ Risk model → coverage matrix, not script soup
→ Four reporters, trace-on-retry, artifact retention
→ k6 load baseline on the critical path
proof: playwright-sdet-regression-suite — 37/37 in CI
A working automation — intake, triage, routing, RAG-backed answers — built on n8n/Make or LangGraph, with human approval points where they belong and logs you can audit.
One visible, production-grade improvement: shipped, measured, in your repo.
03 · Build
~4–8 weeks· scoped & quoted
The full system: suite, gate, automation, runbook, owned by your team at handoff.
04 · Operate
ongoing· optional
Measure, improve, publish: the system compounds instead of quietly decaying.
spec OFFER/01 · fixed price
OFFER/01
AI Quality Audit
$497
About a week, fixed price.
I point my evaluation and QA tooling at your live AI feature, map the highest-leverage failure surface — correctness, safety, the test and automation gaps — and hand you a prioritized plan, the findings that matter, and a real quote for the build you own. The $497 credits straight into the build if you continue. The build itself is quoted from what the audit finds, so you pay for your problem, not a package.
The fixed-price way to start. Everything downstream of it — the build — stays scoped and quoted from what the audit turns up.
✓Every engagement is quoted in writing before work starts: fixed scope, never open-ended hours. The audit fee credits straight into the build if you continue. If scoping shows I can't help, I'll say so and it costs nothing.
one engagement at a time — currently booking Q4 2026
how a typical 4-week engagement runs
WEEK 1
Scope & risk map
Your feature, your failure modes. We agree what "working" means and how it will be measured.
WEEK 2–3
Build
The automation, agent, or suite — in your repo, your CI, your conventions.
WEEK 3–4
Harness & gate
Evals and regression wired into CI. A red gate blocks the deploy, not the retro.
HANDOFF
Runbook & evidence
Docs, scorecard, and a walkthrough so your team owns it without me.
common questions
We already have QA. Why bring you in?
Your QA team almost certainly covers the deterministic surface. The gap is the LLM feature: no golden set, no judge, no gate. Prompt changes ship on vibes. I add the eval layer to your existing pipeline; your team owns it when I leave.
Do you work in our stack or yours?
Yours. Your repo, your CI, your conventions. Every engagement ends with a runbook and a walkthrough. The deliverable is a system your team runs without me. There's no dependency on me.
What does an engagement cost?
Common needs have a fixed price — a $497 AI Quality Audit, packaged builds from $4,997, and monthly care from $299/mo. The entry audit and the monthly retainers have set prices; project work is quoted in writing after a short scoping call — fixed scope, never open-ended hours. The 15-minute intro call is free and ends with a concrete plan and a real number either way. Start with the audit and you get the quote for the larger build baked in.
How fast can you start?
Booking one engagement at a time (see the availability chip at the top of the page). Week 1 is scoping, so the calendar risk is low: you'll know exactly what you're getting before the build starts.
What's the typical timeline?
The audit is one week. Pilot-sized sprints run 2–3 weeks; full builds 4–8 depending on scope. Every engagement ships something visible inside the first two weeks. No long silent phases.
How do you handle our data and code?
NDA-friendly. Access is scoped to the engagement, credentials stay in your systems, code stays in your repos, and nothing is retained after handoff. The data-handling one-pager is available before we start.
Couldn't we just build this in-house?
Often you should — and at handoff you will be. The difference is role: an in-house hire maintains the system day to day; I'm brought in to design it once, with the pattern from many of these builds already baked in — the eval battery, the failure modes, the gate. The build hours are the cheap part. The real cost of doing it in-house is the ongoing one: keeping evals honest, chasing model and API drift, and the regression nobody catches until a customer does. I stand that up so it holds, then hand your team the keys.
After handoff — what if it breaks, and what if you're not around?
Everything lands in your repo and CI, with a runbook and a walkthrough — no dependency on me, nothing locked to my platform. If something drifts, the gate is built to catch it before it ships and the runbook says exactly what to do. Access, credentials, and a written handoff plan exist from day one, so the work outlives me being unavailable. The optional operate window is there if you want a hand, not because you'll need one.
What if the evals show our AI feature is fine?
Then you ship with evidence instead of hope. That's the win. The suite stays in CI catching the regression that would have landed six weeks from now.
Full-time roles?
Open to the right one. The "Hire me" toggle up top reframes this page for recruiters, and the résumé PDF is one click away.
05
05 — about
JTJason Teixeira
For thirteen years my job was making sure software actually worked. The kind of QA where a missed bug costs real money. Then AI features started shipping on a demo and a prayer, and I watched good products break in ways a customer found first. So now I build both halves: the AI, and the proof it works.
— Jason Teixeira · founder, Sage Ideas
I run Sage Ideas, an AI-native studio, and I've spent years building the same thing at every layer: systems that refuse to say "done" without evidence.
That thesis shows up as an engineering OS with a hash-chained proof ledger, a QA platform where every score is computed from a real command, and agent pipelines with golden-trace evals wired into CI. The pattern is always the same: ship the feature, then ship the machinery that catches it when it regresses.
If your team is shipping LLM features on vibes (no eval suite, no regression gate, no idea whether last week's prompt change made things worse), that's exactly the gap I close.
I build
LangGraph · Mastra agents
RAG — pgvector · citations
n8n · Make · Zapier flows
Next.js · FastAPI · Supabase
voice agents — Retell · Vapi
JTone engineer
I prove
Playwright · Pytest · k6
Promptfoo · DeepEval
LLM-as-judge — faithfulness
jailbreak · injection · PII evals
CI gates · signed evidence
capability
typical AI builder
typical QA engineer
me
Ships LLM features (RAG, agents, automation)
✓
—
✓
Builds eval suites & LLM-as-judge gates
rare
rare
✓ shipped
Regression discipline in CI (Playwright/Pytest)
—
✓
✓ 13 yrs
AI-safety battery (injection · PII · toxicity)
—
—
✓ 10 runners
Evidence you can audit (signed, reproducible)
—
rare
✓ by design
Certified for it (ISTQB CT-AI + TAE)
—
rare
✓ both
06
06 — track record
Fortune 50 rigor, founder speed.
Thirteen years across enterprise retail and fintech, plus an AI-native studio. Every role in the same discipline: automation that proves software works. Figures from prior roles are self-reported. Full résumé (PDF) ↓ · Company overview · See the proof →
Built everything in the index above: the agent warehouse, the proof-first engineering OS, and an 85-runner QA platform with ten dedicated LLM-safety evals.
Nexural Research: 58 FastAPI endpoints · 604 tests · 93% coverage, AI response validation that cross-references model claims against deterministic source metrics.
✓Every engagement is quoted in writing before work starts: fixed scope, never open-ended hours. The audit fee credits straight into the build if you continue. If scoping shows I can't help, I'll say so and it costs nothing.
one engagement at a time — currently booking Q4 2026
free mini-eval · a work sample you keep
Send me your AI feature. I’ll run a free mini-eval against it and send you the findings: pass/fail, failure examples, and what I’d gate before your next release. No call required.
1 — intro call
15 minutes, free. Bring the feature that scares you.
2 — risk map
Week 1: we agree what "working" means and how it's measured.
3 — fixed quote
Price attached to deliverables instead of open-ended hours. Either way, the plan is yours.