JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Detect Hallucinations Automatically in Production Logs
How-to guides

Detect Hallucinations Automatically in Production Logs

Flag confident-but-wrong answers in production automatically, by checking each claim against its sources.

A grounding check takes an answer your system already gave, breaks it into claims, and checks each claim against the sources that were supposed to back it. If a claim has no support in the sources, that is your hallucination signal. This guide builds one that runs over your production logs and flags the confident-but-wrong answers. Budget about forty minutes.

#Before you start

  • A RAG or LLM feature that logs the question, the final answer, and the retrieved source text.
  • Python 3.10+ and the OpenAI SDK (pip install openai).
  • An API key in an env var (OPENAI_API_KEY). The examples use it as the judge.

#Log the three things a grounding check needs

You cannot check grounding without the sources that were in context when the answer was made. Log the question, the final answer, and the exact retrieved chunks for every request. One JSONL line per request is enough. If you only log the answer, you have nothing to ground against, so fix the logging first.

logs/answers.jsonl
{"id": "req-1041", "question": "What is the refund window?", "answer": "You can get a full refund within 30 days, no questions asked.", "sources": ["Refunds are available within 14 days of purchase. A restocking fee may apply."]}
{"id": "req-1042", "question": "Do you ship to Canada?", "answer": "Yes, we ship to Canada in 5-7 business days.", "sources": ["We currently ship to the United States and Canada. Delivery takes 5-7 business days."]}

#Write the grounding check

The check asks a model one narrow question: is this answer fully supported by these sources? Keep the judge's job small and force a structured verdict, so you get a label you can act on instead of an essay. Ask for the unsupported claim too, because that is what a human will want to see.

ground.py
import json
from openai import OpenAI

client = OpenAI()

PROMPT = """You are a grounding checker. Decide if the ANSWER is fully supported by the SOURCES.
An answer is grounded only if every factual claim in it appears in or follows directly from the sources.
Contradictions or invented specifics (numbers, dates, policies) are NOT grounded.

QUESTION: {question}
SOURCES: {sources}
ANSWER: {answer}

Reply as JSON: {{"grounded": true|false, "unsupported_claim": "<the worst unsupported claim, or empty>"}}"""

def check(record):
    msg = PROMPT.format(
        question=record["question"],
        sources="\n".join(record["sources"]),
        answer=record["answer"],
    )
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": msg}],
        response_format={"type": "json_object"},
        temperature=0,
    )
    return json.loads(resp.choices[0].message.content)

#Run it over the log file

Loop the check across your JSONL log and print only the answers that failed. This is the whole point: turn a pile of logs into a short list of suspect answers with the exact claim that was not supported. Start on a sample of a few hundred lines before you point it at everything.

scan.py
import json
from ground import check

flagged = 0
total = 0
with open("logs/answers.jsonl") as f:
    for line in f:
        record = json.loads(line)
        total += 1
        verdict = check(record)
        if not verdict["grounded"]:
            flagged += 1
            print(f"[UNGROUNDED] {record['id']}: {verdict['unsupported_claim']}")

print(f"\n{flagged}/{total} answers ungrounded ({flagged/total:.0%})")

#Turn the rate into an alert

A single bad answer is noise. A rising ungrounded rate is a signal worth waking up for. Compute the rate over a window and exit non-zero when it crosses a threshold, so a cron job or CI step can page you or fail the build. Set the threshold from your own baseline. Do not copy a number from here.

gate.py
import sys

MAX_UNGROUNDED_RATE = 0.05  # tune to your baseline

def gate(flagged, total):
    rate = flagged / total if total else 0
    print(f"ungrounded rate: {rate:.1%} (threshold {MAX_UNGROUNDED_RATE:.0%})")
    if rate > MAX_UNGROUNDED_RATE:
        print("FAIL: hallucination rate above threshold")
        sys.exit(1)
    print("OK")

if __name__ == "__main__":
    flagged, total = int(sys.argv[1]), int(sys.argv[2])
    gate(flagged, total)

#Send the flagged answers somewhere a human will see them

The scan is only useful if someone reads the flags. Write the ungrounded answers to a file or a channel your team actually watches, and review them on a schedule. Feed the real hallucinations back into your golden set as known failures. That closes the loop between detection and prevention.

#Watch out for

  • The judge is itself an LLM and can be wrong. It will miss subtle hallucinations and sometimes flag correct answers, so treat the output as a triage queue you still review by hand. Spot-check its calls before you trust the rate.
  • This only catches claims that contradict or overreach the sources. If retrieval pulled the wrong sources in the first place, a grounded-looking answer can still be wrong. Grounding checks the answer against the sources. It does not check whether the sources were right.
  • Running a judge call per log line costs money and time. Sample your logs, run on a schedule instead of inline, and use a cheap model for the check.

#What you built

You have a grounding check that reads your production logs, flags answers that are not supported by their sources, and fails when the ungrounded rate climbs. It is honest about its own limits and cheap to run on a sample. Next, wire the flagged cases into your golden set so today's hallucinations become tomorrow's regression tests.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms