Home / Docs
Documentation
What I build, how it works, and the proof behind it.
The complete, honest reference for working with me: every capability documented, the evaluation method explained, and how engagements actually run. Start with the overview, press ⌘K to search, or jump to what you need.
▸
New here? The fastest way to see the work is a free mini-eval on your live AI feature, or the live eval running in your browser.
What I build
The eval method
Working together
Reference
How-to guides
Build an LLM Eval Gate in CI with PromptfooBuild a Golden Set for LLM Regression TestingAdd a Human-Approval Checkpoint to an AI PipelineWrite Adversarial Probes for Prompt Injection TestingEvaluate a RAG Pipeline End-to-End with RAGASSet Up a Nightly LLM Regression Suite in GitHub ActionsVersion Golden Datasets Like CodeBuild a Tool-Call Accuracy Test Harness for AI AgentsTest Multi-Turn Agent ConversationsAdd Open-Source LLM Observability (Langfuse or Phoenix)Build a Quality Ratchet That Blocks RegressionsRun Playwright Tests Against an AI-Generated UIDetect Hallucinations Automatically in Production LogsBuild a Red-Team Test Suite for a Customer-Facing ChatbotWire an Eval Gate into a GitHub or Vercel Deploy PipelineStructure a Monorepo for an AI Product Plus Eval Suite
Prefer to talk it through?
Ask my AI associate on any page, or book a 15-minute call.