JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Structure a Monorepo for an AI Product Plus Eval Suite
How-to guides

Structure a Monorepo for an AI Product Plus Eval Suite

Lay out an AI product repo so prompts, datasets, and evals have a real home instead of scattering.

You will lay out a monorepo where the app, the prompts, the datasets, and the evals each have their own folder and their own boundaries. The point is that six months from now you can find a prompt, trace it to the cases that test it, and change it without a treasure hunt. This uses pnpm workspaces, which is the least surprising way to do a JS/TS monorepo. Budget about forty minutes to get the skeleton standing.

#Before you start

  • Node.js 18 or newer and pnpm 8 or newer installed.
  • A rough idea of your app stack. The example assumes TypeScript, but the layout is language-agnostic.
  • An eval runner in mind. The example wires Promptfoo, but the structure works with any runner.
  • Git initialized in the repo root.

#Draw the top-level boundaries first

Decide the folders before you write code, because moving them later is annoying. Four homes: the app, the shared prompts, the datasets, and the evals. Keep prompts and datasets as their own packages so both the app and the evals can import the exact same thing. That shared import is the whole trick. The app runs a prompt, the eval grades the same prompt, and neither has its own private copy.

repo tree
my-ai-product/
├── apps/
│   └── web/              # your product
├── packages/
│   ├── prompts/          # prompt templates, versioned
│   └── datasets/         # golden cases, fixtures
├── evals/                # eval configs + runner glue
├── pnpm-workspace.yaml
└── package.json

#Declare the workspace

Tell pnpm which folders are packages. This is what lets apps/web and evals depend on @repo/prompts by name instead of by a fragile relative path. Keep the glob list short and explicit.

pnpm-workspace.yaml
packages:
  - "apps/*"
  - "packages/*"
  - "evals"

#Make prompts a real package

Give the prompts their own package and store each prompt as a versioned template file. Name the package in its package.json so others can depend on it. Now a prompt has one home and a git history you can read. When you change wording, the diff shows up in one file.

packages/prompts/package.json
{
  "name": "@repo/prompts",
  "version": "0.0.0",
  "private": true
}

#Keep datasets next to nothing else

Datasets are just data. Store them as CSV or JSONL in their own package so they do not get tangled up with app code or test frameworks. One file per use case keeps diffs small when you add a case. The app can read these as fixtures, and the evals read the same files, so your test cases and your seed data never drift apart.

packages/datasets/support/cancel-flow.csv
input,__expected
How do I cancel my plan?,llm-rubric: Explains that cancellation lives under Settings > Billing and invents no fee
Can I get a refund?,llm-rubric: States the 14-day refund window and promises nothing beyond it

#Wire the evals to the real prompt and dataset

The evals folder holds the runner config. It points at the prompt file and the dataset by relative path. The template uses {{input}}, which matches the column in the CSV, so each case flows into the prompt. The __expected column tells Promptfoo how to grade the answer. Because it reads the same files the app ships, you grade what actually goes out. Run it from the evals folder. If a case fails, Promptfoo exits non-zero, which is what a CI gate needs.

evals/promptfooconfig.yaml
prompts:
  - file://../packages/prompts/support.txt

providers:
  - openai:gpt-4o-mini

tests: file://../packages/datasets/support/cancel-flow.csv

#Add root scripts so nobody guesses the commands

Put the common tasks in the root package.json. New contributors run one command. This is also what your CI calls, so local and CI stay identical. Keep the API key in an env var. Promptfoo reads OPENAI_API_KEY from the environment, so nothing goes in the config.

package.json
{
  "name": "my-ai-product",
  "private": true,
  "scripts": {
    "dev": "pnpm --filter web dev",
    "eval": "cd evals && npx promptfoo@latest eval"
  }
}

#Watch out for

  • Do not let the app keep its own private copy of a prompt. The moment that happens, the eval grades one thing and the user gets another, and your green suite means nothing. Always read from the shared package.
  • Datasets grow. A folder of 500 mixed cases becomes a swamp. Split by use case early, and keep a short README saying where a new case goes, or people will dump files anywhere.
  • A monorepo is not free. If your app and evals are in different languages, pnpm workspaces will not glue them, and you may want a task runner like Turborepo or plain Make instead. Do not adopt heavy tooling before the repo actually hurts.

#What you built

You have a repo where prompts, datasets, and evals each have a real home, and the evals grade the same files the app ships. That shared boundary is what keeps your quality checks honest as the codebase grows. Next, wire the eval command into CI as a required check so a prompt change that lowers quality stops at the pull request.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms