AI · Early access

llama.cpp hosting on your own server

A fast, dependency-light C/C++ engine for running LLMs on ordinary hardware.

Early access. Join the list and we email you when llama.cpp opens on VPS.

What llama.cpp does

llama.cpp is a lightweight inference engine for running Llama-family and other open LLMs efficiently on CPUs or modest GPUs, using quantized model formats to cut memory use. It underpins many other self-hosted AI tools as their inference backend. It also provides Python and server bindings, so other self-hosted AI tools commonly embed it as their inference engine.

llama.cpp at a glance

License
MIT
Source code
github.com/ggml-org/llama.cpp
Website
llama.app
Runs on
Your own server (VPS)
Good to know
Runs on CPU alone; a GPU speeds it up but is not required, especially with quantized models.
  • Low-resource LLM inference

  • Backend for other AI tools

  • Running quantized models on modest hardware

llama.cpp, connected to the rest of your business

  • On your domain

    llama.cpp answers at an address like llamacpp.yourbusiness.com, with SSL and daily backups switched on.

  • Mail from your address

    llama.cpp sends its emails from your CloudWish business mailbox. Business email

  • AI with a cap you set

    Give llama.cpp its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)

Run llama.cpp on your own server.