AI · Early access
Hugging Face hosting on your own server
A self-hosted inference server for running open models from the Hugging Face hub.
Early access. Join the list and we email you when Hugging Face opens on VPS.
What Hugging Face does
This is a ready-to-deploy server for running Hugging Face models yourself, serving text, vision, or speech models through an API instead of calling a hosted inference endpoint. It's used to keep inference costs and data on infrastructure the user controls. It supports the broad catalog of open models published on the Hugging Face Hub, from small classifiers to large language models.
Hugging Face at a glance
- License
- Apache-2.0
- Source code
- github.com/huggingface/text-generation-inference
- Website
- huggingface.co
- Runs on
- Your own server (VPS)
- Good to know
- Needs a GPU for good speed with most modern models; CPU inference is slow for anything beyond small models.
Self-hosted model inference API
Running open-source ML models privately
Avoiding per-call hosted inference fees
Hugging Face, connected to the rest of your business
On your domain
Hugging Face answers at an address like huggingface.yourbusiness.com, with SSL and daily backups switched on.
Mail from your address
Hugging Face sends its emails from your CloudWish business mailbox. Business email
AI with a cap you set
Give Hugging Face its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)