Langfuse
An open-source platform for tracing, evaluating and managing prompts for LLM applications and AI agents.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About Langfuse
Langfuse records hierarchical traces of every LLM call, tool invocation and retrieval step in an AI application, and ties them to prompt management, datasets, experiments, human annotation queues and evaluations. Teams use production traces to find problems, test fixes and compare cost and latency.
It can be used as a hosted service or run on your own servers from the open-source code. SDKs, a public API and data export to blob storage make it possible to connect it to existing tooling. Version 4 focused on real-time performance.
Strengths
- Tracing, prompt management and evaluation in one tool
- Can be self-hosted or used as a hosted service
- Public API, SDKs and export to blob storage
- Annotation queues for human review
Limitations
- Requires instrumenting your application with its SDK
- Self-hosting means running and maintaining the server yourself
Details
- Pricing
- FreemiumOpen source for self-hosting, with a hosted version that can be started for free.
- License
- Open source, license not stated
- Developer
- Langfuse
- Platforms
- Web, Self-hosted
- How it runs
- Self-hosted, Hosted service
- Best suited for
- Teams running LLM apps or agents in production who need tracing and evals
- Categories
- Developer tools, AI developer tools
- Last verified
- Added
- Sources
Alternatives to Langfuse
Compare allSoftware that can replace Langfuse for an important use case, and what changes if you switch.
LangSmith
A hosted platform for tracing, monitoring and evaluating LLM applications and AI agents.
LangSmith is a proprietary hosted service with detailed agent tracing, SDKs for four languages and PagerDuty alerts, and self-hosting it requires Kubernetes.
Arize Phoenix
A self-hostable tool for tracing, evaluating and experimenting with LLM applications and agents.
Arize Phoenix offers tracing, evals, prompts and experiments that run locally, in Docker or on Kubernetes, under a source-available rather than open-source license.
Opik
An open-source platform from Comet for debugging, evaluating and monitoring LLM applications and agents.
Opik from Comet is open source with tracing, automated evaluations and dashboards, though self-hosting means running several services and contributors must sign a CLA.
MLflow
An open-source platform for tracing, evaluating and managing LLM applications, agents and machine learning models.
MLflow is Apache-2.0 with OpenTelemetry tracing, 50+ evaluation metrics and prompt optimization, and also covers classic model training.
Helicone
An open-source AI gateway and observability platform for routing, logging and debugging LLM requests.
Helicone works as an Apache-2.0 gateway that logs requests and tracks costs across providers, with less prompt and eval tooling, and its roadmap is uncertain after joining Mintlify.
Weights & Biases
A hosted platform for tracking machine learning experiments, evaluating models and tracing LLM applications.
Weights & Biases is a proprietary hosted platform combining experiment tracking and LLM tracing, with no self-hosted option listed.
Langfuse as an alternative
Listings that name Langfuse as an alternative.
DeepEval
An open-source Python framework for unit-testing and evaluating the outputs of large language model applications.
Langfuse pairs evaluations with tracing, prompt management and annotation queues in a hosted or self-hosted platform, requiring you to instrument your application with its SDK.
Evidently
An open-source framework for evaluating, testing and monitoring LLM applications, RAG systems and ML models.
Langfuse centres on tracing, prompt management and evals for LLM apps in a hosted or self-hosted platform, without classic ML drift checks.
Promptfoo
A command-line tool for testing prompts, evaluating model output and red-teaming LLM applications and agents.
Langfuse is an open-source web platform combining tracing, prompt management and evaluation, which requires instrumenting your application with its SDK rather than running tests from a command line.
Similar software
Related functionality, not necessarily a direct replacement.
LiteLLM
A self-hosted AI gateway that calls over 100 LLM providers through one OpenAI-compatible API.
Label Studio
An open-source data labeling tool for building training datasets and reviewing AI output.