AI · Early access

Hugging Face hosting on your own server

A self-hosted inference server for running open models from the Hugging Face hub.

Early access. Join the list and we email you when Hugging Face opens on VPS.

What Hugging Face does

This is a ready-to-deploy server for running Hugging Face models yourself, serving text, vision, or speech models through an API instead of calling a hosted inference endpoint. It's used to keep inference costs and data on infrastructure the user controls. It supports the broad catalog of open models published on the Hugging Face Hub, from small classifiers to large language models.

Hugging Face at a glance

License
Apache-2.0
Source code
github.com/huggingface/text-generation-inference
Website
huggingface.co
Runs on
Your own server (VPS)
Good to know
Needs a GPU for good speed with most modern models; CPU inference is slow for anything beyond small models.
  • Self-hosted model inference API

  • Running open-source ML models privately

  • Avoiding per-call hosted inference fees

Hugging Face, connected to the rest of your business

  • On your domain

    Hugging Face answers at an address like huggingface.yourbusiness.com, with SSL and daily backups switched on.

  • Mail from your address

    Hugging Face sends its emails from your CloudWish business mailbox. Business email

  • AI with a cap you set

    Give Hugging Face its own AI key with a monthly spend cap. Usage shows as its own line on your bill. AI in your apps (early access)

Run Hugging Face on your own server.