lm-evaluation-harness

EleutherAI's command-line framework for few-shot evaluation of language models on many benchmark tasks.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About lm-evaluation-harness

lm-evaluation-harness is a Python tool from EleutherAI for measuring how language models perform on standard benchmarks. It runs few-shot evaluations across hundreds of tasks, and new tasks can be defined with YAML templates.

It is widely used by researchers and model developers who want reproducible scores to compare models or track changes during training and fine-tuning. The project is actively developed on GitHub, with a large contributor base and thousands of commits.

Strengths

  • Hundreds of benchmark tasks in one tool
  • New tasks defined through YAML templates
  • Widely used, so scores are comparable across projects

Limitations

  • Command-line and Python only, no graphical interface
  • Running large benchmarks needs substantial compute

Details

Pricing
FreeFree and open source.
License
MIT
Developer
EleutherAI
Platforms
Windows, macOS, Linux, Command line
How it runs
Downloadable app
Account
Not required
Works offline
Yes
Best suited for
Researchers and developers benchmarking language models
Last verified
Added

Alternatives to lm-evaluation-harness

Compare all

Software that can replace lm-evaluation-harness for an important use case, and what changes if you switch.

  • Promptfoo

    A command-line tool for testing prompts, evaluating model output and red-teaming LLM applications and agents.

    Promptfoo compares prompts and models side by side from one npx command and adds red-teaming, but it tests your application rather than running hundreds of standard benchmarks.

  • DeepEval

    An open-source Python framework for unit-testing and evaluating the outputs of large language model applications.

    DeepEval treats LLM output checks as Python unit tests for applications instead of academic benchmark suites, and judge-based evaluations may need API access to a model.

  • Kiln

    A free desktop workbench for building, evaluating and fine-tuning AI systems, with an open-source library.

    Kiln moves evaluation into a free desktop app with fine-tuning and synthetic data, replacing the command line with a graphical interface, though the app's license is not stated.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing