Build a Red-Team Test Suite for a Customer-Facing Chatbot
Give your chatbot a standing set of attacks it has to survive before every release.
A red-team suite is a fixed set of attacks your chatbot has to survive before you ship it: prompt injection, jailbreaks, data leaks, and off-topic abuse. You run it on every release. If a new answer leaks a system prompt or promises a refund it should not, the build stops. This guide wires one up with Promptfoo, which has a built-in red-team runner, and puts it behind a CI gate. Budget about forty-five minutes.
#Before you start
- Node.js 18 or newer installed.
- An API key for the model your bot uses (the examples use
OPENAI_API_KEY). - A way to call your chatbot: an HTTP endpoint or a prompt you can point Promptfoo at.
- A repo with CI (the gate example uses GitHub Actions).
#Initialize the red-team config
Run the red-team init in your repo. It writes a starter promptfooconfig.yaml with a redteam section and asks what you are testing. This differs from a normal eval: Promptfoo generates the adversarial inputs for you instead of you hand-writing every case.
npx promptfoo@latest redteam init#Point it at your real chatbot
Edit the config so the target is your actual bot. A bare model call skips your guardrails. Use an HTTP provider if your chatbot sits behind an endpoint, so the probes hit the same guardrails your users do. Describe the bot in purpose, because the attack generator uses that to write relevant probes.
targets:
- id: https
config:
url: "https://api.yourapp.com/chat"
method: POST
headers:
Content-Type: application/json
Authorization: "Bearer {{env.CHATBOT_API_KEY}}"
body:
message: "{{prompt}}"
transformResponse: "json.reply"
redteam:
purpose: >
A customer-facing support bot for a SaaS billing product. It should answer
account and billing questions and refuse anything else.
numTests: 5#Choose the attacks that match your risk
Promptfoo groups probes into plugins (what to test for) and strategies (how to deliver the attack). Pick the ones that map to real damage for a customer bot: prompt extraction, PII leaks, hijacking to off-topic tasks, and making promises the business cannot keep. Keep the list short at first so the run stays fast and readable.
redteam:
plugins:
- harmful:misinformation-disinformation
- pii:direct
- prompt-extraction
- hijacking
- excessive-agency
strategies:
- jailbreak
- prompt-injection#Generate and run the probes
Generate the adversarial test cases, then run them. The first command writes the probes to a file so the same attacks run every release instead of being regenerated each time. The eval scores each probe pass or fail from your bot's response. The report opens a readable view of what got through.
npx promptfoo@latest redteam generate -o redteam-tests.yaml
npx promptfoo@latest eval -c redteam-tests.yaml
npx promptfoo@latest redteam report#Gate the release on it
Add a CI job that runs the saved probes on every pull request. promptfoo eval exits non-zero when a probe gets through, which fails the job. Commit the generated redteam-tests.yaml so the suite is fixed and reviewable, and keep both keys in CI secrets.
name: redteam
on: [pull_request]
jobs:
redteam:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
- name: Run red-team suite
run: npx promptfoo@latest eval -c redteam-tests.yaml
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
CHATBOT_API_KEY: ${{ secrets.CHATBOT_API_KEY }}#Watch out for
- Red-team probes are generated by a model and are non-deterministic. Commit the generated test file so every release faces the same attacks. Regenerate when you mean to.
- A green run means the bot survived these attacks. It does not mean the bot is safe. Real attackers are creative and your plugin list is small. Treat a pass as a floor. Add every real incident back as a fixed probe.
- Running probes against a live endpoint sends adversarial traffic through your real stack. Point it at a staging bot, or expect the attempts to show up in your logs, rate limits, and any downstream costs.
#What you built
You have a standing red-team suite that fires a fixed set of injections, jailbreaks, and leak attempts at your chatbot on every pull request, and blocks the merge when one gets through. It is small and honest. It turns \"we think the bot is safe\" into a check that has to pass. Grow the plugin list and add every production incident as a new probe, so the suite gets meaner as your bot meets the real world.