Context recall
Context recall asks a blunt question. When the answer lived in one of your documents, did retrieval actually pull that document back? It measures the fraction of the needed information that made it into the context window. If the right chunk never gets retrieved, nothing downstream can save you. The model is answering from a stack that is missing the page that mattered.
Why it matters
Everything else in a RAG pipeline sits downstream of this. Your prompt can be perfect and your model sharp, but if the chunk with the answer never showed up, the best case is "I don't know" and the worst is a confident guess. Low recall is the quiet failure mode. The system looks like it's working while it simply can't see the evidence it needs. Faithfulness checks whether the answer matches the sources. Context recall checks whether the source was even there to match.
How it works
You need questions paired with the ground-truth passage or facts that answer them. For each one, run retrieval and check whether the retrieved chunks contain that ground-truth information, then average across your test set. Tools like RAGAS score it by breaking the reference answer into claims and counting how many are supported by what got retrieved. When recall is low, the usual suspects are a chunking strategy that slices the answer in half, weak embeddings, or a top-k set too small. Hybrid search often lifts it by catching exact terms that vector search alone misses.
A support bot gets asked "how long do I have to return a laptop?" The answer sits in a returns doc that says 30 days for electronics. But the retriever pulls three chunks about shipping and warranties and never grabs that line. Context recall for this question is zero, and no amount of prompt tuning will fix it until the returns chunk starts coming back.