vLLM

High-throughput, memory-efficient inference and serving engine for LLMs.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About vLLM

vLLM is a library for fast, cost-efficient LLM inference and serving, using techniques like continuous batching to maximize throughput. It was originally developed at UC Berkeley's Sky Computing Lab and is now maintained as an open-source project for self-hosted model serving.

Strengths

  • High-throughput serving suited to production-scale local deployments
  • Apache-2.0 licensed and actively developed

Limitations

  • Aimed at developers deploying models rather than casual end users
  • Requires appropriate GPU hardware for good performance

Details

Pricing
FreeFree and open source under the Apache-2.0 license.
License
Apache-2.0
Developer
The vLLM project
Platforms
Linux, Command line
How it runs
Downloadable app, Self-hosted
Best suited for
High-throughput, memory-efficient inference and serving engine for LLMs
Categories
Local AI tools
Last verified
Added
Provenance
Selected from the TechWalrus Resource Hub (AI); facts checked against the developer's own pages, 3 sources on file.

Report a wrong fact or a dead link on this listing