JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Evaluate a RAG Pipeline End-to-End with RAGAS
How-to guides

Evaluate a RAG Pipeline End-to-End with RAGAS

Measure both halves of a RAG system, so you know whether a wrong answer is a retrieval bug or a generation bug.

RAGAS scores a RAG pipeline on two fronts: how good the retrieval was, and how faithful the answer was to what got retrieved. That split is the point. It tells you whether a wrong answer came from bad context or from the model making things up. This guide runs RAGAS over a small eval set and reads the numbers. Budget about thirty minutes.

#Before you start

  • Python 3.9 or newer.
  • An OpenAI API key in OPENAI_API_KEY. RAGAS uses an LLM to judge and defaults to OpenAI.
  • A RAG pipeline you can call to get an answer and the retrieved chunks for a question.
  • A handful of questions with known correct answers.

#Install RAGAS

Install RAGAS and the datasets helper it uses to hold your eval rows. RAGAS calls an LLM to grade some metrics, so set your key in the environment. Never put the key in code.

terminal
pip install ragas datasets
export OPENAI_API_KEY=sk-...

#Assemble the eval rows

RAGAS needs four fields per row. The question, your pipeline's answer, the contexts your retriever returned as a list of strings, and the ground_truth answer you already know is correct. Run each question through your real pipeline to fill answer and contexts. Do not hand-write the contexts. The whole point is to grade what your retriever actually pulled.

build_dataset.py
from datasets import Dataset

# each contexts entry is the list of chunks your retriever returned
data = {
    "question": [
        "What is the refund window?",
        "How do I cancel my plan?",
    ],
    "answer": [
        "You have 14 days to request a refund.",
        "Go to Settings, then Billing, then Cancel.",
    ],
    "contexts": [
        ["Refunds are available within 14 days of purchase."],
        ["To cancel, open Settings > Billing > Cancel plan."],
    ],
    "ground_truth": [
        "Refunds are available within 14 days.",
        "Cancel from Settings > Billing > Cancel.",
    ],
}

dataset = Dataset.from_dict(data)

#Pick metrics that split retrieval from generation

Choose metrics that map to each half of the system. context_precision and context_recall grade retrieval, whether the right chunks showed up. faithfulness and answer_relevancy grade generation, whether the answer stuck to the context and addressed the question. Keeping both groups is what lets you tell a retrieval bug from a generation bug.

metrics.py
from ragas.metrics import (
    context_precision,
    context_recall,
    faithfulness,
    answer_relevancy,
)

metrics = [
    context_precision,   # retrieval
    context_recall,      # retrieval
    faithfulness,        # generation
    answer_relevancy,    # generation
]

#Run the evaluation

Call evaluate with the dataset and your metrics. RAGAS runs each metric over every row and returns aggregate scores between 0 and 1. This costs a few API calls per row because the LLM-graded metrics ask a model to judge, so keep the set small while you iterate.

run_eval.py
from ragas import evaluate
from build_dataset import dataset
from metrics import metrics

result = evaluate(dataset, metrics=metrics)
print(result)
# {'context_precision': 0.95, 'context_recall': 0.90,
#  'faithfulness': 0.88, 'answer_relevancy': 0.91}

#Read it as retrieval vs generation

Now diagnose. If context precision or recall is low, the retriever is the problem and the generator never had a chance, so fix chunking, embeddings, or top-k. If retrieval scores are high but faithfulness is low, the context was fine and the model invented something, so that is a generation bug. Export the per-row scores to find the exact questions that dragged the average down.

per_row.py
from ragas import evaluate
from build_dataset import dataset
from metrics import metrics

result = evaluate(dataset, metrics=metrics)

# one row per question, with every metric as a column
df = result.to_pandas()
print(df[["question", "context_recall", "faithfulness"]])
df.to_csv("ragas_scores.csv", index=False)

#Watch out for

  • RAGAS grades with an LLM, so scores wobble a little between runs. Treat them as directional. Watch how they move across a fixed eval set instead of reading a single decimal place.
  • context_recall needs ground_truth. If you skip that field the metric silently returns nothing, so give every row a real known-good answer.
  • The default judge is an OpenAI model and it costs money per row. A big eval set on every commit gets expensive fast. Keep the set small, or run the full set nightly and a subset on each change.

#What you built

You now score a RAG pipeline on both halves and can tell whether a bad answer came from retrieval or generation. That split turns vague RAG debugging into a specific fix in a specific component. Next, wire the aggregate scores into a gate so a change that drops faithfulness or recall fails the build.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms