MLflow
An open-source platform for tracing, evaluating and managing LLM applications, agents and machine learning models.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About MLflow
MLflow covers the development loop for AI applications. It captures full traces of LLM applications and agents on top of OpenTelemetry, works with any LLM provider and agent framework, and monitors quality, cost and safety in production.
Evaluation offers more than 50 built-in metrics and LLM judges, or your own. Prompts can be versioned, tested and optimized with lineage tracking, and an AI gateway routes requests across LLM providers. It also tracks classic model training experiments. It suits teams that want an open, self-hostable alternative to proprietary observability tools.
Strengths
- Tracing built on OpenTelemetry, independent of LLM provider
- 50+ built-in evaluation metrics and LLM judges
- Prompt versioning and automatic prompt optimization
- Covers both LLM apps and classic model training
Limitations
- Broad feature set takes time to learn
- Running it for a team means hosting and maintaining a server yourself
Details
- Pricing
- FreeFree and open source.
- License
- Apache-2.0
- Developer
- The MLflow contributors
- Platforms
- Web, Self-hosted, Command line
- How it runs
- Downloadable app, Self-hosted
- Account
- Not required
- Best suited for
- Teams tracking, evaluating and monitoring LLM apps and ML models
- Categories
- AI developer tools
- Last verified
- Added
- Sources
Alternatives to MLflow
Compare allSoftware that can replace MLflow for an important use case, and what changes if you switch.
Weights & Biases
A hosted platform for tracking machine learning experiments, evaluating models and tracing LLM applications.
Weights & Biases covers experiment tracking and LLM tracing as a hosted proprietary service, so you avoid running a server but send data to their platform.
Langfuse
An open-source platform for tracing, evaluating and managing prompts for LLM applications and AI agents.
Langfuse focuses on LLM tracing, prompt management and evaluation, self-hosted or hosted, but does not cover classic model training tracking the way MLflow does.
Opik
An open-source platform from Comet for debugging, evaluating and monitoring LLM applications and agents.
Opik from Comet offers tracing, automated evaluations and dashboards for LLM apps and agents, self-hostable, but it focuses on LLMs rather than classic ML training.
Arize Phoenix
A self-hostable tool for tracing, evaluating and experimenting with LLM applications and agents.
Arize Phoenix handles tracing, evals, prompts and experiments locally, in Docker or on Kubernetes, but uses a source-available license rather than Apache-2.0.
LangSmith
A hosted platform for tracing, monitoring and evaluating LLM applications and AI agents.
LangSmith is a proprietary hosted platform for agent tracing, online evals and alerts, with self-hosting requiring Kubernetes, and it does not track classic model training.
Evidently
An open-source framework for evaluating, testing and monitoring LLM applications, RAG systems and ML models.
Evidently tests and monitors both LLM systems and classic ML models, including drift checks, from Python code under Apache-2.0, without MLflow's tracing server and model management.
TensorBoard
A visualization toolkit for inspecting machine learning training runs, metrics and model graphs.
TensorBoard charts training runs and model graphs locally without an account, but it is oriented toward TensorFlow and lacks LLM tracing and evaluation features.
DeepEval
An open-source Python framework for unit-testing and evaluating the outputs of large language model applications.
DeepEval is a Python framework that treats LLM output checks as unit tests, covering evaluation only, with no tracing server or model registry.
MLflow as an alternative
Listings that name MLflow as an alternative.
Kubeflow
A Kubernetes-native set of open-source projects for running data, ML and AI workloads.
MLflow covers experiment tracking, evaluation and model management with tracing, rather than Kubernetes-native training and serving, and needs a self-hosted server for teams.
Promptfoo
A command-line tool for testing prompts, evaluating model output and red-teaming LLM applications and agents.
MLflow is an Apache 2.0 platform with 50+ built-in evaluation metrics, LLM judges and prompt versioning, and it also covers classic model training, though its broad feature set takes time to learn.
Similar software
Related functionality, not necessarily a direct replacement.
ZenML
An open-source framework for writing ML pipelines and AI agents in Python and running them on your own infrastructure.
Flyte
An open-source orchestration platform for durable data, machine learning and AI agent workflows written in Python.
Helicone
An open-source AI gateway and observability platform for routing, logging and debugging LLM requests.