vLLM
High-throughput, memory-efficient inference and serving engine for LLMs.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About vLLM
vLLM is a library for fast, cost-efficient LLM inference and serving, using techniques like continuous batching to maximize throughput. It was originally developed at UC Berkeley's Sky Computing Lab and is now maintained as an open-source project for self-hosted model serving.
Strengths
- High-throughput serving suited to production-scale local deployments
- Apache-2.0 licensed and actively developed
Limitations
- Aimed at developers deploying models rather than casual end users
- Requires appropriate GPU hardware for good performance
Details
- Pricing
- FreeFree and open source under the Apache-2.0 license.
- License
- Apache-2.0
- Developer
- The vLLM project
- Platforms
- Linux, Command line
- How it runs
- Downloadable app, Self-hosted
- Best suited for
- High-throughput, memory-efficient inference and serving engine for LLMs
- Categories
- Local AI tools
- Last verified
- Added
- Provenance
- Selected from the TechWalrus Resource Hub (AI); facts checked against the developer's own pages, 3 sources on file.