llama.cpp
C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.
Overview
C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.
- Category
- LLM runtimes · AI tools
- Platforms
- Linux, macOS, Windows, API only
- Pricing
- Free
- License
- MIT
- Source code
- ggml-org/llama.cpp
- Repository stars
- 128,722
- API
- No public API description
Release history
Updated 1 day ago10 releases in the last 12 months.
-
b11030
-
b11033
-
b11034
-
b11035
-
b11036
-
b11037
What users say
No reviews yet
Ratings appear once verified reviews arrive. Until then the rank uses public signals only: release freshness and repository adoption.
Alternatives to llama.cpp
See all alternatives Ollama
3
93.8
=
Runtime to download and run large language models locally, with a CLI and a REST API.
Updated 2 days ago
vLLM
22
91.3
=
High-throughput LLM serving engine with paged attention and an OpenAI-compatible server.
Updated 2 days ago
LocalAI
54
89.1
=
Self-hosted, OpenAI-compatible API that runs text, image and audio models on local hardware.
Updated 2 days ago