JTjason.teixeira() Docs
services Book a call →
Home / Docs / How-to guides / Add Open-Source LLM Observability (Langfuse or Phoenix)
How-to guides

Add Open-Source LLM Observability (Langfuse or Phoenix)

See what your LLM app is actually doing in production, with open-source tracing you host yourself.

Observability for an LLM app means you can see each call: the prompt, the model output, the latency, the token cost, and where a chain went wrong. This guide stands up Langfuse on your own machine with Docker, then wires your Python app to send traces to it. Budget about thirty minutes. It also points at Phoenix if you want the OpenTelemetry route instead.

#Before you start

  • Docker and Docker Compose installed.
  • Python 3.9 or newer with an LLM app or a small script to instrument.
  • An API key for your model (the examples use OPENAI_API_KEY).

#Run Langfuse locally with Docker

Clone the repo and start the stack. This brings up the Langfuse server plus its Postgres and other dependencies. Once it is up, open the UI at localhost:3000 and create an account and a project. It is fully self-hosted, so no data leaves your machine.

terminal
git clone https://github.com/langfuse/langfuse.git
cd langfuse
docker compose up

#Get your keys and set them as env vars

In the Langfuse UI, go to Settings and create a set of API keys for your project. You get a public key and a secret key. Put them in your environment along with the host. Keep them out of your code.

terminal
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_HOST="http://localhost:3000"
export OPENAI_API_KEY="sk-..."

pip install langfuse openai

#Trace model calls with the drop-in wrapper

The fastest way to start is the OpenAI drop-in. Change your import so calls route through Langfuse, and every request gets traced automatically with prompt, response, latency, and token counts. Your calling code stays the same.

app.py
# was: import openai
from langfuse.openai import openai

response = openai.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarize observability in one line."}],
)
print(response.choices[0].message.content)

#Group multi-step work into one trace

Real features are more than one model call. Wrap your own functions with the @observe decorator so a retrieval step, a model call, and a post-process show up as nested spans under a single trace. Now you can see the whole flow as one trace.

pipeline.py
from langfuse import observe
from langfuse.openai import openai

@observe()
def retrieve(question: str) -> str:
    # your real retrieval goes here
    return "context about the question"

@observe()
def answer(question: str) -> str:
    context = retrieve(question)
    response = openai.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": f"Use this context: {context}"},
            {"role": "user", "content": question},
        ],
    )
    return response.choices[0].message.content

answer("What is LLM observability?")

#Open the trace and read it

Run your app, then open the project in the Langfuse UI and click into a trace. You will see the nested spans, the exact inputs and outputs, latency per step, and token cost. This is your feedback loop when a call is slow, wrong, or too expensive. Traces send in the background, so in a short-lived script call flush before exit.

flush.py
from langfuse import get_client

# call once before a short script exits so buffered traces are sent
get_client().flush()

#Watch out for

  • Traces are sent asynchronously in a background thread. A short script can exit before they upload, so call flush() at the end or you will see nothing in the UI. Long-running servers handle this on their own.
  • The docker compose stack is meant for local and evaluation use. For a real production deployment you need to set strong secrets, put it behind auth, and back up Postgres. Do not expose the default compose setup to the internet as-is.
  • If you would rather standardize on OpenTelemetry, use Phoenix instead. Run pip install arize-phoenix, start it with phoenix serve, and instrument with the OpenInference libraries. It has the same goal, so pick one and commit.

#What you built

You now have a self-hosted tracing backend and an app that reports every LLM call and multi-step flow into it, with cost and latency attached. That visibility is the thing you reach for the moment a user says the output was wrong and you have no idea why. Next, add scores to your traces so you can grade outputs and spot quality drops over time.

Want this built into your pipeline?
Get a free mini-eval on your live AI feature, or book a call to have it wired in properly.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms