EN·ES·PT
The method · how every engagement actually runs

The Proof-First Method

Most AI work ships on a demo and a prayer — and a customer finds the one thing it gets wrong. Mine ships behind a gate, with the evidence attached. This is the exact operating system that runs my own two live products, applied to your build: four phases, a concrete artifact produced at each, and the standard enforced before anything moves forward.

Fixed scope, in writing An artifact at every phase Done = verified, not claimed
The loop · four phases, each with a deliverable you keep
Phase 01 · ~1 week

Audit — find where it breaks before you build

What you get

A prioritized plan, a risk map of where your AI can fail, a baseline eval run against your current feature, and a firm, fixed quote for the build — yours to keep whether or not you continue.

The standard enforced

Every claim is measured, nothing is assumed. You leave with a real number and a real plan, not a range and a maybe.

Phase 02 · ~2 weeks

Sprint — ship a thin slice, gated from day one

What you get

A working thin slice of the feature, the first eval harness wired to your real traffic patterns, and a CI gate that already blocks a bad change from merging.

The standard enforced

It ships behind a gate from the first commit. No feature lands without a test that can fail — proof is built in, not bolted on.

Phase 03 · ~4–8 weeks

Build — the feature, and the evidence it works

What you get

The full feature, the eval battery — faithfulness, prompt-injection, regression, safety — a signed proof ledger, and a runbook your team can operate without me.

The standard enforced

"Done" means verified by a command you can re-run, not asserted in a status meeting. No fake green — not even mine.

Phase 04 · optional retainer

Operate — keep it proven in production

What you get

Monitoring, scheduled eval runs, maintained CI quality gates, and a monthly proof report in language your PM can read.

The standard enforced

A regression is caught before your customer sees it — and the result is published, not claimed. You can hand the repo to anyone and it still runs.

The standards · the same ones that govern my own products

Five rules I don't break.

No fake green — not even mine

The gate goes red when it should, in public. This very site publishes its own failing runs; your project gets the same honesty.

Done means verified, not claimed

Every number is backed by an artifact you can open — a run, a ledger, a test exit code. If it can't be shown, it isn't done.

Your repo, your infra, your keys

Work runs on your accounts, not mine. Nothing is retained after handoff — no lock-in, no dependency on me to keep it running.

Priced to the problem, not a package

Fixed scope, quoted in writing before work starts. Never open-ended hours, never a menu price padded for the worst case.

Founder-led, senior-only

The person who scopes it is the person who builds and operates it. No account managers, no junior handoff, no telephone game.

Not a slide · the machine behind the method is real and public

The same system runs my products and your build.

This isn't a methodology I wrote for a sales page. It's the tooling I run every day — and most of it is open for you to inspect.

▸ sage-kernel — the proof-first engineering OS ▸ llm-eval-gate — the open eval template ▸ this site's own QA, in public ▸ the body of work it produced ▸ the quality standard & SLA

It starts with one low-risk week.

The audit is the front door: a plan and a real quote you own, whether or not you build.

Start with the audit → See it proven first