Build an LLM Eval Gate in CI with Promptfoo
Wire a real quality gate into CI so a bad prompt change fails the build instead of reaching your users.
An eval gate is a check that runs your LLM through a set of known cases and fails the build if quality drops. Once it is in place, a bad prompt change stops at the pull request instead of reaching a user. This guide wires one up with Promptfoo, an open-source runner, and GitHub Actions. Budget about thirty minutes.
#Before you start
- Node.js 18 or newer installed.
- An API key for whatever model you use (the examples use
OPENAI_API_KEY). - A repo with GitHub Actions enabled.
#Initialize Promptfoo
Run the init command in your repo. It drops a starter config you will edit next. There is nothing to install globally, npx fetches it for the run.
npx promptfoo@latest init#Write your first test cases
Replace the starter config with real cases. Each test has the input variables and one or more asserts. Mix cheap deterministic checks (contains) with an llm-rubric for the nuanced part. Start with five to ten cases that cover a common path and a known failure.
prompts:
- "You are a support assistant. Answer the question: {{question}}"
providers:
- openai:gpt-4o-mini
tests:
- vars:
question: "How do I cancel my plan?"
assert:
- type: contains
value: "Settings"
- type: llm-rubric
value: "Explains the real cancellation steps and does not invent a cancellation fee"
- vars:
question: "Do you offer refunds?"
assert:
- type: llm-rubric
value: "States the 14-day refund window and does not promise anything beyond it"#Run it locally
Run the eval from the terminal and open the viewer to see each case pass or fail with the model's actual output. This is your feedback loop while you tune prompts.
npx promptfoo@latest eval
npx promptfoo@latest view#Wire it into GitHub Actions
Add a workflow that runs the eval on every pull request. Promptfoo exits non-zero when an assert fails, and a non-zero exit fails the job. That failing job is your gate. Put your API key in the repo secrets, never in the file.
name: eval-gate
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
- name: Run eval gate
run: npx promptfoo@latest eval -c promptfooconfig.yaml
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}#Make it a required check
In the repo's branch-protection settings, mark the eval-gate job as a required status check for merging. Now a pull request that lowers quality cannot be merged until someone fixes it or consciously updates the cases. That last part matters: the gate should be hard to bypass by accident and easy to change on purpose.
#Watch out for
- LLM-graded asserts are non-deterministic. Keep the rubric tight and specific, and do not set the bar so high that a good answer fails half the time.
- Cost adds up if the suite is huge and runs on every commit. Start small, and consider running the full suite nightly and a fast subset on each pull request.
- A gate that everyone routinely overrides is theatre. If people keep bypassing it, the cases are wrong, not the people. Fix the cases.
#What you built
You now have a check that runs real cases against your model on every pull request and blocks the merge when quality drops. It is small, it is honest, and it is the single highest-leverage thing you can add to an LLM feature. Grow the case set from production over time, especially from the answers that embarrassed you.