AI · Early access
llama.cpp hosting on your own server
A fast, dependency-light C/C++ engine for running LLMs on ordinary hardware.
Early access. Join the list and we email you when llama.cpp opens on VPS.
What llama.cpp does
llama.cpp is a lightweight inference engine for running Llama-family and other open LLMs efficiently on CPUs or modest GPUs, using quantized model formats to cut memory use. It underpins many other self-hosted AI tools as their inference backend. It also provides Python and server bindings, so other self-hosted AI tools commonly embed it as their inference engine.
llama.cpp at a glance
- License
- MIT
- Source code
- github.com/ggml-org/llama.cpp
- Website
- llama.app
- Runs on
- Your own server (VPS)
- Good to know
- Runs on CPU alone; a GPU speeds it up but is not required, especially with quantized models.
Low-resource LLM inference
Backend for other AI tools
Running quantized models on modest hardware
llama.cpp, connected to the rest of your business
On your domain
llama.cpp answers at an address like llamacpp.yourbusiness.com, with SSL and daily backups switched on.
Mail from your address
llama.cpp sends its emails from your CloudWish business mailbox. Business email
AI with a cap you set
Give llama.cpp its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)