AI · Early access

vLLM hosting on your own server

A high-throughput serving engine for running large language models.

Early access. Join the list and we email you when vLLM opens on VPS.

What vLLM does

vLLM is an inference server optimized for serving LLMs to many concurrent users efficiently, using techniques like paged attention to maximize GPU utilization. Teams self-host it to run open models at production-grade throughput instead of relying on a hosted API. It supports most popular open model families out of the box, including Llama, Mistral, and Qwen.

vLLM at a glance

License
Apache-2.0
Source code
github.com/vllm-project/vllm
Website
vllm.ai
Runs on
Your own server (VPS)
Good to know
Needs a GPU for good speed; the whole point of vLLM is maximizing GPU throughput under load.
  • High-throughput LLM API serving

  • Production self-hosted model inference

  • Batched multi-user LLM workloads

vLLM, connected to the rest of your business

  • On your domain

    vLLM answers at an address like vllm.yourbusiness.com, with SSL and daily backups switched on.

  • Mail from your address

    vLLM sends its emails from your CloudWish business mailbox. Business email

  • AI with a cap you set

    Give vLLM its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)

Run vLLM on your own server.