Sage Ideas. Jason Teixeira · AI engineering & quality
Statement of Work
Ref: SOW-SAMPLE-01
Version 1.0 · Illustrative
agency.sageideas.dev
Statement of Work · build plan

LLM Feature Evaluation & Quality Gate

A representative engagement: take one AI feature you're not yet confident shipping, build the evaluation and CI gate that keeps it honest, and hand you a repeatable quality system your team can run without me. This document shows exactly how I scope, build, and prove the work.

Sample document. This is an illustrative build plan, not a binding quote. Real scope, timeline, and price are confirmed after a short call and written into a project-specific SOW. Figures below are indicative ranges, not fixed prices.
1

Parties & engagement

ProviderSage Ideas LLC — Jason Teixeira (hello@sageideas.dev)
Client[Client legal name]
Engagement typeFixed-scope project · one AI feature · your repository
Estimated duration~4 weeks from kickoff (see §5)
Working modelRemote · async-first with a weekly checkpoint call · your stack, your git
2

Objective & success criteria

The problem this solves. Most teams ship an AI feature and then hope it behaves. "It seems to work" is not a release criterion — and it fails quietly, in front of customers, the first time a prompt, a model version, or a retrieval index drifts.

What "done" means here. The engagement is successful when every one of these is true and demonstrable:

3

Scope of work

Delivered in four phases, each ending in something you can see and verify — not a status update.

01Discovery & risk map~week 1
Review the feature, its prompts, data flow, and failure modes. Agree the quality dimensions that matter (accuracy, grounding, safety, latency, cost) and the thresholds that define "shippable." Produce a one-page risk model so coverage tracks real risk, not guesswork.
DeliverableRisk model + evaluation plan (agreed thresholds)
02Build the evaluation harness~weeks 1–2
Author the golden set from real and adversarial inputs. Implement scoring (deterministic checks + LLM-as-judge where judgment is needed), the safety battery, and per-run quality/latency/cost measurement. Every check is version-controlled in your repo.
DeliverableGolden set + scoring harness + safety battery, in your repo
03Wire the CI gate~week 3
Connect the harness to your CI so a pull request that regresses quality, safety, or cost is blocked with a readable report. Add a ratchet so the bar only moves up. Prove it works by making it fail on purpose, then pass.
DeliverablePassing CI gate + a captured red→green run proving it fires
04Handoff & enablement~week 4
Write the runbook: how to add a case, read a failure, tune a threshold, and interpret the gate. Walk your team through it live. The goal is that the system outlives the engagement and needs no consultant to operate.
DeliverableRunbook + recorded walkthrough + handoff session
4

Deliverables summary

DeliverableFormAcceptance
Risk model & evaluation planDocument, agreed thresholdsSigned off in the week-1 checkpoint
Golden evaluation setVersioned data in your repoCovers the agreed dimensions
Scoring harness + safety batteryCode in your repoRuns locally & in CI, green
CI quality gateCI workflowBlocks a seeded regression; passes clean
Runbook + walkthroughDocs + recordingYour team runs it unaided
5

Timeline

WeekFocusCheckpoint
1Discovery, risk map, evaluation plan; start the golden setPlan sign-off
2Harness, scoring, safety batteryHarness demo
3CI gate + ratchet; prove it firesRed→green captured
4Runbook, enablement, handoffFinal acceptance

Timeline assumes timely access (repo, CI, a representative environment) and one client point of contact for decisions. Slippage in access shifts the schedule, not the scope.

6

Out of scope

To keep the engagement fixed and honest, the following are explicitly excluded unless added in writing: building the underlying feature itself; model training or fine-tuning; production infrastructure changes; ongoing on-call or monitoring (available separately as a retainer); and evaluation of features beyond the one named in §1.

7

Assumptions & client responsibilities

8

Commercials

This engagement is quoted as a fixed price after a scoping call — you approve the number before any work starts, in writing. Indicative range for a single-feature evaluation & gate:

ComponentModelIndicative
AI Quality Audit (entry point)Fixed · credits into the build$497
Full evaluation & CI gate buildFixed-scope project$8k–$18k
Ongoing quality retainerMonthly · optional$2.5k–$4k / mo

Indicative ranges for illustration only; not an offer. Deposit of 30% books the engagement; balance on final acceptance. Exact figures are set in the project SOW after a call.

9

Next step

One 15-minute call turns this sample into a real, fixed quote for your feature — with a plan you keep either way. agency.sageideas.dev/book.html
Jason Teixeira
Sage Ideas LLC · Provider
Signature & date
[Client signatory]
[Client legal name] · Client
Signature & date

This sample is provided for illustration and does not itself create a binding agreement. A project-specific Statement of Work, executed under a Master Services Agreement, governs any engagement.

Sample SOW — download PDF ↓ · ← portfolio