EN·ES·PT
Case studies · proof you can open

Four systems. Four measured outcomes. Every one links to the receipt.

No logos I can't back up and no numbers I can't show you. Each of these is a system I built or rescued, with the metric that mattered, and a link to the public repo, the verbatim run, or the live screenshot that proves it.

Fortune-50 test engineering ISTQB Test Automation Engineer Open-source: llm-eval-gate · playwright-sdet This site runs its own QA — 100+ checks, axe-clean
Outcomes · what the numbers meant for the business
15

exploitable vulnerabilities blocked from shipping — in public, by the gate, before a customer or auditor ever saw them.

10% → <1%

a flaky suite the team had learned to ignore, turned back into a release gate they trust.

4h → 75m

release regression cut from half a day — safer, faster ships across systems serving 2,300+ stores.

0

uncited answers in a customer-facing AI — zero fabricated-source liability, by construction.

Each figure traces to a receipt below. Prior-role numbers are self-reported and labeled as such.

Built & operated · my own products, not client logos

Before the track record: two live products I built and run by myself.

The strongest proof isn't a testimonial — it's software I ship and operate solo, every day. Two products on one engineering system. Open the front door of either.

Flagship 1 · trading-media platform

Nexural

A live SaaS platform — member cockpit, admin ops, Stripe billing, real market-data engines, and a companion Discord AI engine. Production scale, one operator.

~800API routes
86scheduled jobs
244RLS tables
▸ full case study
Flagship 2 · learning platform

Sage Ideas Academy

A live learning product — AI tutor on a real model, interactive labs, Stripe billing, and cryptographically verifiable certificates with a public verify page.

30+courses
400+lessons
~180quality gates
▸ full case study
Track record · my own products & prior roles — each links to a receipt
AI quality & evaluation · nexural-qa-os

The QA platform that refused to show a fake green.

red → green A release carrying 15 exploitable CVEs — caught by the gate at 11:39, blocked, and green again by 14:35. Both runs published verbatim, so “passing” is a fact you can audit, not a claim.
The problem

A quality platform is only trustworthy if it can fail. On a real run, the dependency-audit gate found 15 high/critical CVEs. The honest verdict was not shippable, not a green badge.

What I did
  • Let the gate go red and published the verbatim run: no quiet override
  • Fixed the drift at the source: 9 pinned security floors, killing the unpatchable transitive chain
  • Re-ran the full battery: 85 runners incl. 10 LLM-safety evals
The result

3,759 tests · 91% coverage · 13/13 gates. High/critical CVEs: 15 → 0. The whole caught→blocked→patched→proven arc is on the site, both captures included.

▸ the red run (verbatim) ▸ the green rerun verified 2026-08-15 · both runs published
Test automation · playwright-sdet-regression-suite

A regression suite a release manager can trust.

37/37 specs green in CI · 15.3s · 0 flakes — a regression estate the team can trust to block a bad merge, with traces, screenshots, and a written risk model.
The problem

Most suites are script soup: coverage follows whoever wrote tests last, red is ignored because it's flaky, and every failure starts a screenshot scavenger hunt.

What I did
  • Risk model → coverage matrix, so coverage tracks release risk
  • Page Object Model, fixtures, trace-on-retry, four reporters
  • Wired into CI with artifact retention, a green badge that means something
The result

37/37 specs, zero flakes, 15.3s. The evidence folder — traces and screenshots — ships in the public repo. Clone it and reproduce the run.

▸ public repo ▸ the run verified · run dated 2026-07-10, evidence in repo
AI engineering · RAG research dashboard

An AI assistant that can't answer without a citation.

100% of answers cite their source — built into the design, not requested by prompt — so nothing a customer reads can invent a source. There is no uncited generation path, and no fabricated-source liability.
The problem

Most RAG demos hallucinate confidently. "Grounded" is claimed in the prompt and quietly violated in production the first time retrieval comes back thin.

What I did
  • Built retrieval so every answer is bound to ranked, selected chunks
  • Removed the uncited generation path entirely: abstain when there's no evidence
  • Tracked quality, safety, latency, and cost per query on an analytics endpoint
The result

100% citation coverage, by construction. Verified live. I ran it, asked a real question, and captured the actual cited answer below.

The RAG research dashboard answering a real query with chunk-level citations: ranked retrieval candidates selected, 100% citation coverage, and an honest no-evidence abstain rate in the metrics row.
▸ full-size screenshot verified 2026-08-15 · ran live, real query captured
Prior role · HighStrike · 2021–2026

Making red mean something again on a live suite.

10% → <1% flake rate on a live production suite — cut with retry logic and test-isolation fixes, no rebuild required. From a prior full-time role; figures self-reported.
The problem

A flaky suite is worse than none: when red is noise, the team learns to ignore it, and the real regression sails through with the false alarms.

What I did
  • Diagnosed the flake sources: shared state, timing, ordering
  • Installed a flake protocol: quarantine lane, retry policy, isolation fixes
  • Kept the existing suite and treated rebuild as the last resort
The result

Flake rate 10% → under 1%. Red became a real signal again, carried on the same Fortune-50-grade discipline behind 500+-test infrastructure earlier in my career.

▸ the service this became ▸ full track record prior full-time role · figures self-reported
Selected builds & systems · public code is linked, products I run are described

Not a portfolio of mockups. Real systems you can open — shipped, and a few in active build.

Fifty-plus original repositories and a dozen live products across AI/LLM evaluation, QA & test automation, cloud infrastructure, security, and quantitative systems. A representative slice — every public repo links to real code; forks and other people's libraries are deliberately excluded.

Flagships
Open source & infrastructure
sage-kernel★ 1

A proof-first, MCP-native engineering OS for the terminal — 124 MCP tools where every "done" is backed by a command you can re-run.

JavaScript · MCPGitHub ↗
llm-eval-gate

Your first green LLM eval gate in ten minutes — zero API keys to start. The eval-harness pattern I bring to client LLM features, MIT-licensed.

JavaScript · MITGitHub ↗
playwright-sdet-regression-suite

A Playwright + TypeScript SDET regression suite with release-style QA evidence — traces, screenshots, and a reproducible run.

TypeScript · PlaywrightGitHub ↗
terraform-aws-modules

Four production AWS modules — multi-AZ VPC, S3 + CloudFront static site, Lambda API, and GitHub OIDC keyless CI/CD. CI-tested, with variable validation.

Terraform · AWSGitHub ↗
TOGAF Documentation Template★ 2

The full TOGAF ADM enterprise-architecture deliverable map — every document you should produce across all ADM phases, in one checklist.

Enterprise architectureGitHub ↗
micro-saas-starter★ 1

A production SaaS starter: Next.js 16, Supabase auth, Stripe billing — the boilerplate I clone to ship a monetised app fast.

Next.js · Supabase · StripeGitHub ↗
Master Migration Pipeline★ 1

An enterprise migration pipeline with automated deployment and rollback strategies — the runbook I use to move production systems safely.

CI/CD · infraGitHub ↗
graphify

An AI coding-assistant skill that turns any input into a knowledge graph — works across Claude Code, Codex, Cursor, and Gemini.

Python · AI skillGitHub ↗
Products I build & operate
Nexural

A live trading-media platform — ~800 API routes, 244 RLS-secured tables, real market data, a Discord AI engine, and a full learning academy. Built and operated solo.

Next.js · Supabase · AICase study →
Sage Ideas Academy

An AI-engineering learning platform — dozens of courses and hundreds of lessons with in-browser Pyodide labs, a RAG-grounded tutor, and verifiable certificates.

Next.js · Supabase · RAGCase study →
Voza

A speaking-first English-learning app for Latin America — a 256-screen Expo product with a four-tier verification pipeline.

Expo · React Nativeoperated · private
Knox

An AI car-diagnostic app — a mobile-first flow that turns symptoms and photos into a ranked diagnosis.

AI · mobilein build · private
Undeny

An AI insurance-denial appeal builder — turns a denial letter into a grounded, citation-backed appeal.

AI · LLMin build · private
Hard Things Daily

A daily-discipline product — a web + mobile + API monorepo with a production certification gate.

Next.js · Expo · APIoperated · private
Trayd

An AI receptionist + estimator for trades — answers, qualifies, and produces a photo-based estimate.

AI · voicein build · private
Owly

A speed-reading training app — eight mini-games and a character-driven progression loop.

TypeScript · mobileoperated · private

Fifty-plus original repositories in total. Browse them all on GitHub ↗

Want your system to be the next one on this page?

Start with a free mini-eval on your live AI feature, or book a call and we'll scope the fastest path to a measured outcome.

Book a call → or get the sample eval report →

Want senior work below market? I'm taking a few founding engagements at a reduced rate in exchange for a public, named case study.