llama.cpp

C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.

by ggml.ai No reviews yet Open source Free b11030 Updated 1 day ago 9 92.6 =

Overview

C/C++ inference engine for LLMs on CPUs and GPUs, using the GGUF format and quantization.

Category
LLM runtimes · AI tools
Platforms
Linux, macOS, Windows, API only
Pricing
Free
License
MIT
Source code
ggml-org/llama.cpp
Repository stars
128,722
API
No public API description

Release history

Updated 1 day ago

10 releases in the last 12 months.

  1. b11030
  2. b11033
  3. b11034
  4. b11035
  5. b11036
  6. b11037

What users say

No reviews yet

Ratings appear once verified reviews arrive. Until then the rank uses public signals only: release freshness and repository adoption.

Alternatives to llama.cpp

See all alternatives