hyperlakeDiscuss a deployment โ†—
Answer ยท Self-hosting LLMs

How do you deploy an LLM on premises?

Seven steps from choosing a model to running a monitored, access-controlled endpoint, with a go-live checklist.

Updated
The short answer

To deploy an LLM on premises, pick an open-weight model whose licence fits your use, size GPU memory for its weights and KV cache, serve it with an engine such as vLLM, put an authenticated gateway in front, restrict the network, monitor latency and usage, and plan how models and software get updated without internet access.

An on-premises LLM is a model you run on servers you control, reached through an endpoint only your users and systems can call. Downloading the model is the easy part. Most of the work is sizing, access control and keeping the service healthy.

What are the steps to deploy an LLM on premises?

There are seven, and each one feeds the next.

  1. Write down the workload. Who will call the model, how many requests a day, how long the prompts and answers are, the longest context you need, the response time users will accept and the most sensitive data it will see. These numbers drive every later choice.
  2. Choose a model and read its licence. Shortlist two or three open-weight models and test them on your own tasks, not only on public benchmarks. Licences differ: the Llama 3.1 licence, for example, requires a "Built with Llama" notice and a separate licence from Meta above 700 million monthly active users, while others, such as Mistral 7B Instruct v0.3, use Apache-2.0.
  3. Size the hardware. GPU memory must hold the model weights, the KV cache for every request in flight and some overhead. An 8-billion-parameter model at 16-bit precision needs about 16 GB for its weights alone. The answer on how many GPUs you need works through the formula for an 8B and a 70B model.
  4. Install a serving engine. The engine loads the model, batches requests and exposes an API. vLLM, for example, offers OpenAI-compatible endpoints for chat completions, completions and embeddings. On Kubernetes, KServe can run vLLM behind a standard inference resource.
  5. Put a gateway in front. Users and applications should never call the engine directly. A gateway signs callers in through your identity provider, issues per-team keys, applies rate limits and logs every call. vLLM's own --api-key option only covers paths under /v1, /v2 and /inference, and its documentation warns against relying on it alone.
  6. Lock down the network and the supply chain. Allow only the gateway to reach the engine, block outbound traffic the service does not need, and pull weights and images from an internal registry with checksums. The OWASP Top 10 for LLM Applications 2025 lists the risks to test for, including prompt injection, sensitive information disclosure, supply chain and unbounded consumption.
  7. Monitor, load test and plan updates. vLLM publishes Prometheus-format metrics at /metrics, including vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds, vllm:kv_cache_usage_perc and vllm:num_requests_waiting. Load test with realistic prompts before go-live, and keep the previous model version ready to roll back to.

Which serving engine should you use?

Use an engine built for many concurrent users for a shared service; local runtimes suit development and single users.

Engine Fits Note
vLLM Shared GPU serving with an OpenAI-compatible API Prometheus metrics and a built-in benchmark tool
SGLang High-throughput GPU serving Compared with TensorRT-LLM in this post
NVIDIA TensorRT-LLM, often behind Triton Inference Server NVIDIA GPU fleets tuned per model See the post on Triton deployment and dynamic batching
llama.cpp or Ollama One developer's machine, CPUs or small GPUs Simple to start; designed for running models locally
Hugging Face TGI Existing installations only Hugging Face put TGI in maintenance mode and recommends vLLM, SGLang, llama.cpp or MLX

What should you check before go-live?

Run through this list with the people who will operate the service.

  • The model's licence has been read and its conditions recorded.
  • The model was evaluated on your own prompts, and the results are saved.
  • GPU memory leaves room for your peak concurrency at your longest context.
  • A load test at expected peak met your time-to-first-token target.
  • Every call goes through the gateway, signed in through your identity provider.
  • The engine cannot be reached directly from user networks.
  • Outbound traffic is blocked except to approved destinations.
  • Prompts and outputs are logged, or deliberately not logged, according to your data rules, with a retention period.
  • Dashboards and alerts cover latency, queue length, KV cache use, errors and GPU health.
  • Weights, images and packages come from an internal registry, and the upgrade and rollback steps have been rehearsed.
  • Someone is on call for the service.

When should you not deploy on premises?

When volume is low or uncertain, when nobody can run GPU servers, or when the model you need is only available as a hosted service. A hosted API, or a model in your own cloud account, is then often simpler. The answer on whether self-hosting is worth it shows how to find the break-even volume with the cost calculator.

Where does Hyperlake fit?

Hyperlake serves open models through KServe on a Kubernetes foundation and puts one identity path in front of models and data: sign-in through an OAuth/OIDC proxy connected to your identity provider, validated JWTs and OPA policy decisions. It runs in AWS, Azure, Google Cloud, OVH, private cloud or on premises. The FAQ lists what is included.

Frequently asked questions

Can an on-premises LLM run with no internet access?

Yes, if everything it needs is copied in first: model weights, container images, Python packages and drivers. Hugging Face libraries, for example, make no calls to the Hub when HF_HUB_OFFLINE=1 is set and use only cached files. A fully disconnected design is covered in the post on air-gapped AI.

Do I need Kubernetes to run an LLM on premises?

No. One GPU server running a serving engine in a container is enough for a single model. Kubernetes, with an inference layer such as KServe, starts to pay off when you run several models, share GPUs between teams or need rolling upgrades.

Can existing OpenAI SDK code call a self-hosted model?

Usually, yes. vLLM implements OpenAI-compatible endpoints such as /v1/chat/completions, /v1/completions and /v1/embeddings, so many applications only need a new base URL, a key and the new model name. Check any provider-specific features your code relies on.

Is an on-premises LLM automatically compliant with data rules?

No. It removes one external processor, but you still need access control, logging, retention rules and a record of what data the model sees. Compliance depends on how the system is run, not only where.

Sources

Sources for the facts on this page, last checked October 9, 2026.

  1. OpenAI-compatible server (vLLM documentation) checked October 9, 2026
  2. Production metrics (vLLM documentation) checked October 9, 2026
  3. Text Generation Inference: maintenance-mode notice (Hugging Face) checked October 9, 2026
  4. KServe: generative and predictive AI inference on Kubernetes checked October 9, 2026
  5. OWASP Top 10 for LLM Applications 2025 checked October 9, 2026
  6. Environment variables, including HF_HUB_OFFLINE (Hugging Face Hub documentation) checked October 9, 2026
  7. Mistral-7B-Instruct-v0.3 model card, Apache-2.0 licence (Hugging Face) checked October 9, 2026
  8. Llama 3.1 Community License Agreement (Meta) checked October 9, 2026

Start with a workload. Build the environment around it.

Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.