LLM Evaluation & Quality
How to measure whether an LLM feature is good enough to ship — metrics, judges, golden sets, and the evidence that a change did not quietly regress.
All 19 in LLM Evaluation & Quality →A working reference on evaluating, testing, and shipping AI features — organized into eight pillars, from LLM evaluation to workflow automation. Built from real engagements, and every claim links to a receipt. New material lands here continuously; the docs cover how engagements run.
How to measure whether an LLM feature is good enough to ship — metrics, judges, golden sets, and the evidence that a change did not quietly regress.
All 19 in LLM Evaluation & Quality →Evaluating retrieval-augmented generation end to end: faithfulness, context precision and recall, chunking, reranking, and citation accuracy.
All 9 in RAG & Retrieval Quality →Making multi-step, tool-using agents trustworthy: trajectory evaluation, tool-call accuracy, multi-turn testing, and the non-determinism problem.
Wiring evaluation into the pipeline — eval gates, quality ratchets, regression suites, canary releases, and "no fake green" as an actual workflow.
Classic software quality engineering, sharpened for AI-powered products: flaky tests, Playwright, visual regression, coverage that means something.
All 9 in Test Automation & QA →The product- and leadership-level view: launch checklists, "passing tests isn't proof," build-vs-buy, cost, rollback plans, and hiring for AI quality.
All 10 in Shipping AI Safely →Business-process automation with AI in the loop: support triage, lead qualification, human-approval steps, failure modes, and honest ROI.
All 9 in AI Workflow Automation →Straight comparisons and landscape maps across the eval, RAG, and testing tool ecosystems. Written from hands-on use, not SEO farming.