LangSmith vs Langfuse
If you run LLM apps in production and need to see traces, run evals, and debug bad outputs, you are probably weighing these two. LangSmith is a commercial, hosted platform from the LangChain team. Langfuse is open source and self-hostable, with a hosted option too. The real split is about control and where your data lives.
| Dimension | LangSmith | Langfuse |
|---|---|---|
| What it is | Hosted LLM tracing and eval platform | Open-source LLM observability and evals |
| License | Commercial / hosted | Open source (open-core) |
| Hosting | SaaS, with self-host on higher tiers | Self-host or use their cloud |
| Framework fit | Tightest with LangChain and LangGraph | Framework-agnostic, SDKs plus OpenTelemetry |
| Core features | Tracing, datasets, evals, prompt management | Tracing, datasets, evals, prompt management |
| Data control | Lives on their platform by default | Fully yours when self-hosted |
| Best for | Teams deep in the LangChain stack | Teams who want to own the stack |
Pick LangSmith if you are already building on LangChain or LangGraph and want observability that fits with almost no glue. It also makes sense if you would rather pay for a managed service than run your own infrastructure, and having traces on a vendor platform is fine for your data rules.
Pick Langfuse if you want to self-host so prompt and trace data never leaves your infrastructure, or if you are not tied to LangChain. It suits teams who value owning the deployment, a permissive license, and OpenTelemetry-based instrumentation that works across frameworks.
For most teams I reach for Langfuse first. Self-hosting keeps sensitive prompt data in-house, and it does not lock you to one framework. LangSmith is the better call if your whole app is LangChain or LangGraph, since the integration is tight and you skip running your own service. The common mistake is defaulting to LangSmith because you used LangChain for one tutorial, then finding out later you cannot easily keep trace data on your own servers. Decide on data ownership and framework lock-in first. The tracing and eval features are close enough that they rarely settle it.