vLLM
High-throughput LLM serving engine with paged attention and an OpenAI-compatible server.
Overview
High-throughput LLM serving engine with paged attention and an OpenAI-compatible server.
- Category
- LLM runtimes · AI tools
- Platforms
- Linux, API only
- Pricing
- Free
- License
- Apache-2.0
- Source code
- vllm-project/vllm
- Repository stars
- 92,120
- API
- No public API description
Release history
Updated 2 days ago4 releases in the last 12 months.
-
proto-v0.3.0
-
proto-v0.2.0
-
proto-v0.1.0
-
v0.29.0
What users say
No reviews yet
Ratings appear once verified reviews arrive. Until then the rank uses public signals only: release freshness and repository adoption.
Alternatives to vLLM
See all alternatives Ollama
3
93.8
=
Runtime to download and run large language models locally, with a CLI and a REST API.
Updated 2 days ago
llama.cpp
9
92.6
=
C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.
Updated 1 day ago
LocalAI
54
89.1
=
Self-hosted, OpenAI-compatible API that runs text, image and audio models on local hardware.
Updated 2 days ago