EN·ES·PT
Portfolio · 2026 edition · built & proven
AI automation × QA / LLM eval engineer N 28.5384° W 81.3789°

I ship AI features. Then I prove they work.

I'm who you bring in when you're about to ship an LLM feature that could embarrass you in front of a customer. I build it, and the gate that catches it the moment it drifts. see why →

LLM apps and agents, plus workflow automation (RAG, LangGraph, n8n/Make), paired with the regression and eval harness that keeps them honest in production: Playwright, Pytest, k6, Promptfoo, DeepEval, LLM-as-judge. Most people build the AI or test it. I do both, and every claim below links to an artifact.

0/13 proof gates · red→green today
0 tests passing · 0% line coverage
verified · reproducible by one command
one engagement at a time — currently booking Q4 2026
Book a call → Free mini-eval of your AI feature → Résumé (PDF) ↓ GitHub ↗

Quoted in writing before work starts. The audit fee credits into the build. No fit, no charge.

fig. 01: a feature fails eval → it goes back. that's the whole point.
field receipts · in production

Everyone can build an AI demo now. Almost nobody can prove it works in production.

That gap is the whole job, and it's the part I have receipts for. Four, each linked to the verbatim run.

10% → <1%
flake rate cut on a live
production suite · HighStrike
the case study →
37/37
specs green in CI · 0 flakes
15.3s · public, reproducible
the run →
100%
of RAG answers cite a source
by construction, not by prompt
live screenshot →
15 → 0
high/critical CVEs — my own
gate caught them, on my repo
the red run →
01
01 · proof surface
0
tests passing · verified
0%
line coverage · ratcheted
0/13
proof gates · drift caught & fixed today
nexural-qa-os · proof-loop.mjs · two verbatim runs · 2026-08-15

the sales assistant on this site gets red-teamed too. see the 10/10 results →

LangGraphMastran8nMakeZapierpgvectorSupabaseFastAPINext.jsPlaywrightPytestk6PromptfooDeepEvalLLM-as-judgeLangfuseGitHub Actions
LangGraphMastran8nMakeZapierpgvectorSupabaseFastAPINext.jsPlaywrightPytestk6PromptfooDeepEvalLLM-as-judgeLangfuseGitHub Actions
the transformation · what this does for you

Your operation today, into a system you can prove.

Scroll it. The same five-stage flow, from manual-and-unprovable to automated-and-eval-gated. The shape of what I build for you.

02
02 — the index

Five systems, one thesis.

01
Agents · LLM appsESTIMATED

sage-agents

Agent warehouse monorepo: voice agents, outbound SDR, RAG support chat, and workflow orchestrators on Mastra + LangGraph. It ships with a multi-provider LLM router, cost tracking per run and Langfuse tracing, plus a hard TCPA consent gate on every outbound call.

blank page → staged client agent in days instead of weeks (estimated from template scaffold time)
MastraLangGraphVercel AI SDKSupabaseLangfuse
Private · walkthrough on request
intakeweb · phoneLLM routercost-trackedvoiceTCPA gateSDRRAG chatLangfuseevery run traced
~/sage-agents
$ pnpm tsx infra/scripts/create-agent.ts \
inbound-voice acme-inbound-voice
scaffolded @sage/agent-acme-inbound-voice
✓ prompts · tools · evals/golden.json ready
fig. 02 · architecture · live output
02
Autonomous workflowVERIFIED

sage-kernel

Proof-first MCP engineering OS: 140 tools an AI agent drives through policy, signed approvals, and a hash-chained proof ledger. Nothing is "done" because a model said so. A claim-firewall rejects unproven success language.

140 MCP tools · 78 release gates · verified: gate suite + hash-chained proof ledger in the public repo
Node 22MCPSAST + taintSQLite/Postgres
AI agentclaude · cursorpolicysigned approvals140 toolsMCPproof ledgerhash-chained
~/sage-kernel
$ npm run mcp:smoke
passed, 140 tools
$ npm run release:check
✓ 78/78 gates · exit 0
fig. 03 · architecture · live output
03
RAGVERIFIED

ai-research-dashboard

RAG knowledge base where every answer is extractive and cited. Durable ingestion worker, pgvector retrieval, persisted eval runs, and a full audit trail from source to answer.

100% citation coverage · verified 2026-08-15: ran it live, asked a real question, screenshot below is the actual answer
FastAPIpgvectorGeminiReact
Private · walkthrough on request
sourcesdurable ingestchunkerembedpgvectorcosine top-kcited answeraudit trail
The real AI Research Dashboard answering a query with chunk-level citations: retrieval candidates ranked and selected, 100% citation coverage, honest no-evidence abstain rate visible in the metrics row.
fig. 04 · architecture · real UI, real query, 2026-08-15 · click to zoom
04
QA framework + CIVERIFIED

playwright-sdet-regression-suite

Release-critical e-commerce flows under regression with the evidence a release manager would ask for: traces, screenshots, four reporters, and a written risk model. CI uploads artifacts on every push.

37/37 specs · 15.3s · 0 flakes · verified: evidence/ folder in the public repo, run dated 2026-07-10
PlaywrightTypeScriptPOMGitHub Actions
push / PRCIGitHub Actions37 specs4 workers · POMgreen gatetraces · 4 reporters
~/playwright-sdet-regression-suite
$ npx playwright test
Running 37 tests using 8 workers
37 passed (15.3s)
artifacts → evidence/ · traces · JUnit · screenshots
fig. 05 · architecture · live output
05
LLM-eval suiteVERIFIED

nexural-qa-os

85 quality runners under one CLI, including hallucination, jailbreak, prompt-injection, toxicity, and PII-leak evals for LLM features. Every score is computed from a real command and packaged as ed25519-signed evidence.

3,759 tests · 91% coverage · 13/13 · verified 2026-08-15: the CVE gate went red at 11:39, blocked release, and was green by 14:35. Both runs published.
Turbo + pnpmvitestDAG orchestratored25519
red run ↗green rerun ↗Private · walkthrough on request
qa runone CLIDAGplugin runners85 runners10 AI-safetyverdicted25519 signed
~/nexural-qa-os
11:39 ✗ 15 high/critical CVEs → NOT PROVEN 12/13
+9 security floors · pnpm install
14:35 ✓ 0 high/critical → PROVEN 13/13
# caught → blocked → patched → proven. same day.
fig. 06 · architecture · live output
03
03 · engineering briefs

How the work actually went, as spec sheets you can audit.

Outcome first; the full eight-section spec — including the parts most case studies hide — is one click deeper. Every headline number states how it was measured.

spec BRIEF/01 · rev A
BRIEF/01

Feedback triage that runs itself

Make + Gemini workflow

Negative feedback surfaces in seconds now. It used to wait for the Friday review.

Google FormsMake webhookGemini 2.5 Flash · classify + summarizeSheets lognegative? → email alert
fig. 07 · system flow, left to right
Problem

Customer feedback arrived through a form and sat unread. Negative signals surfaced days late, after the customer was already gone.

Measured results

Negative feedback surfaces in real time instead of at end-of-week review — measured as webhook-to-alert latency of seconds vs a manual weekly pass.

view full spec: 6 more sections
Why not a chatbot

Nobody wants to converse with their own feedback queue. The value is deterministic: classify, summarize, route, alert. A chat interface would add latency and remove auditability.

Architecture

Make scenario, four steps end to end (see fig. 1). Each step re-runnable in isolation.

Retrieval / memory

None needed. Each message is classified independently. Deliberately avoided: statelessness keeps the pipeline debuggable.

Human approval points

The AI never replies to a customer. It routes to a human with a summary; the alert email is the handoff. It doesn't resolve anything.

Production safeguards

Structured output schema on the LLM step, fallback label on parse failure, and the raw message always logged alongside the classification.

What I would harden next

A golden set of labeled messages run nightly through Promptfoo to catch classification drift when the model version changes.

stack · Make · Gemini 2.5 Flash · Google Forms · Sheets · Gmail
spec BRIEF/02 · rev A
BRIEF/02

RAG that cites or shuts up

ai-research-dashboard

100% of answers cite their source. The pipeline has no uncited path.

sourcechunkerembed → pgvectorcosine retrievalextractive answer + citationsaudit trail
fig. 08 · system flow, left to right
Problem

Research teams need answers from their own corpus. A RAG system that paraphrases confidently without sources is a liability rather than a tool.

Measured results

100% of answers citation-backed — verified by design: there is no uncited generation path. Quality, safety, latency, and cost tracked per query on the analytics endpoint.

view full spec: 6 more sections
Why not a chatbot

A chat wrapper over the corpus was the obvious build. The actual need was an operations dashboard: sources, ingestion jobs, eval runs, and feedback all inspectable. The Q&A is one feature inside it.

Architecture

FastAPI API + durable ingestion worker feeding fig. 1; provider gateway swaps local deterministic embeddings for Gemini per environment. React dashboard on top.

Retrieval / memory

Persisted queries, answers, and citations. Deleting a source cascades: chunks, embeddings, jobs, payloads, citations. No orphaned memory.

Human approval points

Answer-level feedback endpoint feeds a review loop; eval runs are triggered and persisted by API so a human signs off before provider or threshold changes ship.

Production safeguards

Extractive-only answers (cite or abstain), full audit trail on source/query/eval/feedback events, deterministic local providers for reproducible dev, Gemini fallback-to-local on missing keys.

What I would harden next

Promote eval thresholds into deployment gates so a retrieval regression blocks the release before it reaches the retro.

stack · FastAPI · pgvector · SQLite/Postgres · Gemini · React
spec BRIEF/03 · rev A
BRIEF/03

The QA OS that can’t lie

nexural-qa-os

3,759 tests. On publish day the gate caught real CVE drift, blocked me, and was green again by afternoon.

qa init · detect stackDAG orchestrator85 runners · incl. 10 LLM-safety evalscomputed verdicted25519-signed evidence
fig. 09 · system flow, left to right
Problem

Quality dashboards state numbers they can’t defend. For LLM features it’s worse: teams ship prompt changes with no regression signal at all.

Measured results

3,759 tests, 91.12% line coverage, 13/13 gates. On 2026-08-15 the CVE gate went honestly red (15 high/critical in deps) and blocked its own release — patched and PROVEN again the same afternoon, both runs published verbatim. One command regenerates the scorecard.

view full spec: 6 more sections
Why not a chatbot

An "AI QA assistant" that opines on quality is exactly the failure mode. The verdict must be computed by code from command exit codes. An LLM never grades its own homework here.

Architecture

One CLI driving fig. 1: every runner is a plugin (detect → plan → run → collectEvidence); one hung runner can’t stall a run. Two gates: <15ms pre-commit, full-repo merge gate.

Retrieval / memory

Evidence ledger instead of memory: every run’s exact command, exit code, and parsed metric recorded to a proof ledger, re-verifiable offline.

Human approval points

The autonomous fix loop is bounded. It fixes only what the harness can verify, commits on improvement, reverts on regression, stops on an honest stall for human review.

Production safeguards

Anti-hallucination contract: no score without a backing artifact. Ratcheting coverage floors, honest-skip discipline, ed25519-signed redacted evidence.

LLM-eval coverage

Ten dedicated AI-safety runners: bias, consistency, hallucination, jailbreak, prompt-injection, refusal, toxicity, PII-leak, cost, latency. That's the harness I bring to client LLM features.

stack · TypeScript · Turbo/pnpm · vitest · Playwright · k6 · ed25519
04
04 — work with me

Packaged engagements

Fixed scope, fixed deliverables, evidence included. Each offer is a productized version of something in the briefs above.

Hiring full-time? These are the three capabilities I'd bring to your team on day one, each backed by a shipped system above.

LLM feature QA & eval harness

Your LLM feature gets a regression suite: golden traces, LLM-as-judge scoring for faithfulness and safety, and a CI gate that blocks the deploy when quality drops.

Promptfoo / DeepEval suite on your real traffic patterns
Hallucination, injection & toxicity runners
CI gate + scorecard your PM can read
proof: nexural-qa-os · 85 runners incl. 10 AI-safety evals
Full details →

Test automation + CI setup

A risk-scoped Playwright/Pytest suite with the evidence discipline of a release manager: traces, reports, artifacts, wired into GitHub Actions from day one.

Risk model → coverage matrix instead of script soup
Four reporters, trace-on-retry, artifact retention
k6 load baseline on the critical path
proof: playwright-sdet-regression-suite · 37/37 in CI
Full details →

AI workflow automation build

A working automation (intake, triage, routing, RAG-backed answers) built on n8n/Make or LangGraph, with human approval points where they belong and logs you can audit.

n8n / Make / Zapier or code-level LangGraph
Human-in-the-loop gates, structured outputs
Runbook + handoff so your team owns it
proof: sage-agents templates · feedback-triage pipeline
Full details →
See the full service matrix: every capability, priced & visual
the engagement path · how we work together
01 · Audit
~1 week · start here

One highest-leverage bottleneck, mapped. You leave with a prioritized plan and a concrete quote you own.

Start a conversation →
02 · Sprint
~2 weeks · scoped & quoted

One visible, production-grade improvement: shipped, measured, in your repo.

03 · Build
~4–8 weeks · scoped & quoted

The full system: suite, gate, automation, runbook, owned by your team at handoff.

04 · Operate
ongoing · optional

Measure, improve, publish: the system compounds instead of quietly decaying.

spec OFFER/01 · fixed price
OFFER/01

The Eval Audit

$2,500

One week, fixed price.

I run your live AI feature through a real evaluation battery — correctness, safety, hallucination, prompt-injection — and hand you a scored report, the failing transcripts, and a prioritized fix list. The $2,500 credits straight into the build if you continue. The build itself is quoted from what the audit finds, so you still pay for your problem, not a package.

The one standardized, fixed-price deliverable I sell. Everything downstream of it — the build — stays scoped and quoted from what the audit turns up.
Every engagement is quoted in writing before work starts: fixed scope, never open-ended hours. The audit fee credits straight into the build if you continue. If scoping shows I can't help, I'll say so and it costs nothing.
one engagement at a time — currently booking Q4 2026
how a typical 4-week engagement runs
WEEK 1
Scope & risk map
Your feature, your failure modes. We agree what "working" means and how it will be measured.
WEEK 2–3
Build
The automation, agent, or suite: in your repo, your CI, your conventions.
WEEK 3–4
Harness & gate
Evals and regression wired into CI. A red gate blocks the deploy before it reaches the retro.
HANDOFF
Runbook & evidence
Docs, scorecard, and a walkthrough so your team owns it without me.
common questions
We already have QA. Why bring you in?

Your QA team almost certainly covers the deterministic surface. The gap is the LLM feature: no golden set, no judge, no gate. Prompt changes ship on vibes. I add the eval layer to your existing pipeline; your team owns it when I leave.

Do you work in our stack or yours?

Yours. Your repo, your CI, your conventions. Every engagement ends with a runbook and a walkthrough. The deliverable is a system your team runs without me. There's no dependency on me.

What does an engagement cost?

There's no fixed price list, because scope varies too much for one to be honest. Every engagement is quoted in writing after a short scoping call: fixed scope, never open-ended hours. The 30-minute intro call is free and ends with a concrete plan and a real number either way. Start with the audit and you get the quote for the larger build baked in.

How fast can you start?

Booking one engagement at a time (see the availability chip at the top of the page). Week 1 is scoping, so the calendar risk is low: you'll know exactly what you're getting before the build starts.

What's the typical timeline?

The audit is one week. Pilot-sized sprints run 2–3 weeks; full builds 4–8 depending on scope. Every engagement ships something visible inside the first two weeks. No long silent phases.

How do you handle our data and code?

NDA-friendly. Access is scoped to the engagement, credentials stay in your systems, code stays in your repos, and nothing is retained after handoff. The data-handling one-pager is available before we start.

What if the evals show our AI feature is fine?

Then you ship with evidence instead of hope. That's the win. The suite stays in CI catching the regression that would have landed six weeks from now.

Full-time roles?

Open to the right one. The "Hire me" toggle up top reframes this page for recruiters, and the résumé PDF is one click away.

05
05 — about
JT Jason Teixeira
For thirteen years my job was making sure software actually worked. The kind of QA where a missed bug costs real money. Then AI features started shipping on a demo and a prayer, and I watched good products break in ways a customer found first. So now I build both halves: the AI, and the proof it works.

— Jason Teixeira · founder, Sage Ideas

I run Sage Ideas, an AI-native studio, and I've spent years building the same thing at every layer: systems that refuse to say "done" without evidence.

That thesis shows up as an engineering OS with a hash-chained proof ledger, a QA platform where every score is computed from a real command, and agent pipelines with golden-trace evals wired into CI. The pattern is always the same: ship the feature, then ship the machinery that catches it when it regresses.

If your team is shipping LLM features on vibes (no eval suite, no regression gate, no idea whether last week's prompt change made things worse), that's exactly the gap I close.

I build
LangGraph · Mastra agents
RAG — pgvector · citations
n8n · Make · Zapier flows
Next.js · FastAPI · Supabase
voice agents — Retell · Vapi
JT one
engineer
I prove
Playwright · Pytest · k6
Promptfoo · DeepEval
LLM-as-judge — faithfulness
jailbreak · injection · PII evals
CI gates · signed evidence
capabilitytypical AI buildertypical QA engineerme
Ships LLM features (RAG, agents, automation)
Builds eval suites & LLM-as-judge gatesrarerare✓ shipped
Regression discipline in CI (Playwright/Pytest)✓ 13 yrs
AI-safety battery (injection · PII · toxicity)✓ 10 runners
Evidence you can audit (signed, reproducible)rare✓ by design
Certified for it (ISTQB CT-AI + TAE)rare✓ both
06
06 — track record

Fortune 50 rigor, founder speed.

Thirteen years across enterprise retail and fintech, plus an AI-native studio. Every role in the same discipline: automation that proves software works. Full résumé (PDF) ↓

Sage Ideas · Nexural

Founder, AI Automation & QA Engineer 2024 — present · Orlando, FL / remote
  • Built everything in the index above: the agent warehouse, the proof-first engineering OS, and an 85-runner QA platform with ten dedicated LLM-safety evals.
  • Nexural Research: 58 FastAPI endpoints · 604 tests · 93% coverage, AI response validation that cross-references model claims against deterministic source metrics.

HighStrike

Quantitative Finance & Automation, Senior Analyst / Full-Stack 2021 — 2026 · fintech, remote
  • Kubernetes-based test infrastructure running 500+ tests against live trading workflows ($100k+ daily volume); flaky rate cut 10% → <1% with intelligent retry.
  • Test execution 45 min → 8 min (−82%) via Docker containerization and parallel runs across microservices.
test execution 45m → 8m

The Home Depot

Software Tester → Automation Engineer → Cloud Engineer 2012 — 2021 · Fortune 50
  • Selenium + Python framework for systems serving 2,300+ stores: regression 4 h → 75 min (−70%), 300+ automated cases at 99.5% stability.
  • Led enterprise AWS migration (Terraform, Kubernetes): −35% infrastructure cost, deployments 4 h → 15 min via Jenkins/GitLab pipelines.
regression 4h → 75m
deploys 4h → 15m
credentials
ISTQB CT-AI — AI testing ISTQB Test Automation Engineer ISTQB CTFL AWS Cloud Practitioner AWS ML Foundations TripleTen AI Automation
B.S. Computer Science, Full Sail · B.S. Finance, Kean · English + Portuguese native · Spanish professional
07 — by the numbers
13+
years in quality & automation
140
public repos on GitHub
85
quality runners built
78
release gates, all green
08
08 — field notes

Working notes from the proof-first trenches.

2026·08Eighteen agents audited a live curriculum. One lesson was teaching a false error.Six auditors executing every code claim, six rewriters, six independent verifiers. 73 defects, 17 critical or high. Every one anchored to a verbatim quote. Thirteen minutes of wall clock.5 min · read →2026·08Five pages, five agents, one design file: zero merge conflictsA parallel fleet shipped five production pages in five minutes of wall clock. The speed is the demo; the three rules that made it shippable are the product.4 min · read →2026·08The day my own quality gate blocked meAt 11:39 the proof loop refused the PROVEN verdict. 15 CVEs had drifted into prod deps. Green by 14:35. Both runs published verbatim.4 min · read →2026·08LLM regression testing with Promptfoo in CI: the minimum viable gate30 golden traces, one judge, one failing exit code. The smallest setup that stops a bad prompt change, runnable this afternoon.6 min · read →2026·08No fake green: what a proof ledger taught me about AI agentsAgents will happily report success they never earned. A better prompt wouldn’t fix this. The fix is an evidence gate the agent physically cannot talk its way past.6 min · read →2026·07Your LLM feature needs a regression suite more than a better promptPrompt tweaks feel like progress because nobody is measuring. Golden traces + LLM-as-judge in CI turn "it seems better" into a diff you can gate on.5 min · read →2026·06Automation that never talks to your customerThe best-converting AI workflow I’ve shipped sends zero AI-written messages. Where to put the human approval point, and why it beats full autonomy.4 min · read →

longer-form writing lives at sageafterdark.com · free artifact: the LLM pre-launch eval checklist →

Let's make your AI provable.

30 minutes. Bring the feature that scares you. Leave with a concrete eval-and-automation plan.

One hire, two disciplines.

The engineer who ships the LLM feature and the harness that guards it. Let's talk about your team.

Book an intro call → or just email me Email me
Every engagement is quoted in writing before work starts: fixed scope, never open-ended hours. The audit fee credits straight into the build if you continue. If scoping shows I can't help, I'll say so and it costs nothing.
one engagement at a time — currently booking Q4 2026
free mini-eval · a work sample you keep

Send me your AI feature. I’ll run a free mini-eval against it and send you the findings: pass/fail, failure examples, and what I’d gate before your next release. No call required.

use the form below; put the feature URL in the message and pick any stage

I reply within one business day myself. There's no automated sequence.

1 — intro call

30 minutes, free. Bring the feature that scares you.

2 — risk map

Week 1: we agree what "working" means and how it's measured.

3 — fixed quote

Price attached to deliverables instead of open-ended hours. Either way, the plan is yours.

github/JasonTeixeira linkedin sageideas.dev
TEIXEIRA
copied