Stack & integrations
The whole promise of this work is that it lands in your repo, your CI, and your conventions — so the honest answer to “does it work with our stack?” is almost always yes. Here’s the concrete version.
Everywhere on this site I say your repo, your CI, your conventions. This page makes that specific: what I demonstrably work in, and — just as important — why the parts that matter aren’t locked to any one vendor. The rule underneath all of it: the deliverables are plain, portable artifacts, so they drop into what you already run rather than forcing a migration.
#Version control & CI/CD
The eval gate and the regression suites are the two things that have to live in your pipeline. Both ship as ordinary commands — npm run eval, npx playwright test — so they wire into any runner that can execute a shell step. GitHub Actions is the default (it’s what the sample eval-gate workflow is written for), but nothing about the gate assumes it.
| System | Notes |
|---|---|
| GitHub Actions | The default. The eval gate and Playwright suites ship as ready-to-run workflows, with the scorecard uploaded as a build artifact — see AI evaluation & quality. |
| GitLab CI · CircleCI · Jenkins · Buildkite | The gate is a single command that exits non-zero below the floor, so it’s a few lines in any of these. I port the workflow to yours rather than asking you to adopt mine. |
| Self-hosted / air-gapped runners | Runs on your own runners with no dependency on a hosted service — the suite and its fixtures live in your repo and execute wherever your CI does. |
#LLM providers
The application and eval layers talk to a model behind a typed boundary, so the provider is a configuration choice, not an architecture decision. I pick it per engagement on capability, cost, latency — and data policy — rather than defaulting to one vendor.
| Provider | Notes |
|---|---|
| OpenAI | GPT models — function-calling and schema-constrained structured outputs, the same typed-boundary pattern shown in AI product engineering. |
| Anthropic (Claude) | Long-context reasoning and tool use; a common choice where grounded, careful answers matter more than raw throughput. |
| DeepSeek | Strong capability-per-dollar — a sensible default when a workload is cost-sensitive and the quality bar is met under eval. |
| Google Gemini | Used where its multimodal or long-context strengths fit the workload. |
#Automation platforms
For intake, triage, and routing workflows the platform is a means, not a religion — I reach for the lightest tool that fits the complexity and who has to maintain it. The short version:
| Reach for | When | Trade-off |
|---|---|---|
| Make / Zapier | Simple, linear glue between SaaS apps you already pay for | Fast to build; brittle past a few branches |
| n8n (self-host) | Branching logic, your own data, or you want to own the runtime | More control; you host and maintain it |
| Code (LangGraph / typed) | AI in the loop, real state, retries, approval gates, tests | Most robust and testable; needs an engineer to change |
#The default stack, where there’s no constraint
When a build starts greenfield with no existing stack to honor, this is what I reach for — chosen for speed to production and low maintenance, not novelty. It’s the same default documented in Product & platform, and every piece is swappable for your equivalent:
Frontend Next.js (App Router) · React · TypeScript
Backend Serverless functions · typed APIs · Zod validation
Data Supabase / Postgres · row-level security
Payments Stripe
Quality Playwright + axe in CI · self-hosted fonts · strict CSP
Automation n8n / Make / LangGraph where a workflow fits