AI · Early access
vLLM hosting on your own server
A high-throughput serving engine for running large language models.
Early access. Join the list and we email you when vLLM opens on VPS.
What vLLM does
vLLM is an inference server optimized for serving LLMs to many concurrent users efficiently, using techniques like paged attention to maximize GPU utilization. Teams self-host it to run open models at production-grade throughput instead of relying on a hosted API. It supports most popular open model families out of the box, including Llama, Mistral, and Qwen.
vLLM at a glance
- License
- Apache-2.0
- Source code
- github.com/vllm-project/vllm
- Website
- vllm.ai
- Runs on
- Your own server (VPS)
- Good to know
- Needs a GPU for good speed; the whole point of vLLM is maximizing GPU throughput under load.
High-throughput LLM API serving
Production self-hosted model inference
Batched multi-user LLM workloads
vLLM, connected to the rest of your business
On your domain
vLLM answers at an address like vllm.yourbusiness.com, with SSL and daily backups switched on.
Mail from your address
vLLM sends its emails from your CloudWish business mailbox. Business email
AI with a cap you set
Give vLLM its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)