Add Open-Source LLM Observability (Langfuse or Phoenix)
See what your LLM app is actually doing in production, with open-source tracing you host yourself.
Observability for an LLM app means you can see each call: the prompt, the model output, the latency, the token cost, and where a chain went wrong. This guide stands up Langfuse on your own machine with Docker, then wires your Python app to send traces to it. Budget about thirty minutes. It also points at Phoenix if you want the OpenTelemetry route instead.
#Before you start
- Docker and Docker Compose installed.
- Python 3.9 or newer with an LLM app or a small script to instrument.
- An API key for your model (the examples use OPENAI_API_KEY).
#Run Langfuse locally with Docker
Clone the repo and start the stack. This brings up the Langfuse server plus its Postgres and other dependencies. Once it is up, open the UI at localhost:3000 and create an account and a project. It is fully self-hosted, so no data leaves your machine.
git clone https://github.com/langfuse/langfuse.git
cd langfuse
docker compose up#Get your keys and set them as env vars
In the Langfuse UI, go to Settings and create a set of API keys for your project. You get a public key and a secret key. Put them in your environment along with the host. Keep them out of your code.
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_HOST="http://localhost:3000"
export OPENAI_API_KEY="sk-..."
pip install langfuse openai#Trace model calls with the drop-in wrapper
The fastest way to start is the OpenAI drop-in. Change your import so calls route through Langfuse, and every request gets traced automatically with prompt, response, latency, and token counts. Your calling code stays the same.
# was: import openai
from langfuse.openai import openai
response = openai.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize observability in one line."}],
)
print(response.choices[0].message.content)#Group multi-step work into one trace
Real features are more than one model call. Wrap your own functions with the @observe decorator so a retrieval step, a model call, and a post-process show up as nested spans under a single trace. Now you can see the whole flow as one trace.
from langfuse import observe
from langfuse.openai import openai
@observe()
def retrieve(question: str) -> str:
# your real retrieval goes here
return "context about the question"
@observe()
def answer(question: str) -> str:
context = retrieve(question)
response = openai.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": f"Use this context: {context}"},
{"role": "user", "content": question},
],
)
return response.choices[0].message.content
answer("What is LLM observability?")#Open the trace and read it
Run your app, then open the project in the Langfuse UI and click into a trace. You will see the nested spans, the exact inputs and outputs, latency per step, and token cost. This is your feedback loop when a call is slow, wrong, or too expensive. Traces send in the background, so in a short-lived script call flush before exit.
from langfuse import get_client
# call once before a short script exits so buffered traces are sent
get_client().flush()#Watch out for
- Traces are sent asynchronously in a background thread. A short script can exit before they upload, so call flush() at the end or you will see nothing in the UI. Long-running servers handle this on their own.
- The docker compose stack is meant for local and evaluation use. For a real production deployment you need to set strong secrets, put it behind auth, and back up Postgres. Do not expose the default compose setup to the internet as-is.
- If you would rather standardize on OpenTelemetry, use Phoenix instead. Run pip install arize-phoenix, start it with phoenix serve, and instrument with the OpenInference libraries. It has the same goal, so pick one and commit.
#What you built
You now have a self-hosted tracing backend and an app that reports every LLM call and multi-step flow into it, with cost and latency attached. That visibility is the thing you reach for the moment a user says the output was wrong and you have no idea why. Next, add scores to your traces so you can grade outputs and spot quality drops over time.