hyperlakeDiscuss a deployment ↗
Answer · Self-hosting LLMs

How many GPUs do I need to serve an LLM?

The memory formula, worked through for Llama 3.1 8B and 70B at a stated context length and concurrency, and what changes the answer.

Updated
The short answer

The number of GPUs to serve an LLM starts with memory: model weights, KV cache for every token in flight, and up to 20% overhead. For 16 requests of 8,192 tokens at 16-bit, Llama 3.1 8B needs about 37 GB (one 48 GB GPU) and 70B about 212 GB (four 80 GB GPUs). Throughput sets how many copies you need.

Two numbers decide the GPU count. Memory sets the minimum: the model and its working state must fit. Throughput sets the rest: enough copies of the model to serve your traffic at the response time you want. This page works out the first and shows where the second comes from.

What is the formula for LLM GPU memory?

GPU memory needed ≈ 1.2 × weights + KV cache, kept under about 92% of the GPUs' total memory.

Term Formula Source
Weights parameters × bytes per parameter EleutherAI, 2023
Overhead up to 20% of the weights, for activations and the runtime EleutherAI's rule of thumb: total ≈ 1.2 × model memory
KV cache per token 2 × layers × KV heads × head size × bytes per value NVIDIA, 2023
KV cache KV cache per token × concurrent requests × tokens per request NVIDIA, 2023
Usable share vLLM lets an instance use 0.92 of GPU memory by default vLLM engine arguments

Bytes per parameter: 4 at 32-bit, 2 at 16-bit (BF16 or FP16), 1 at 8-bit (FP8 or INT8), about 0.5 at 4-bit. The leading 2 in the KV formula is one key and one value per layer. NVIDIA's version multiplies by attention heads; models with grouped-query attention store only their KV heads, which is why the formula above uses KV heads.

How much memory does Llama 3.1 8B need?

About 37 GB at 16-bit for 16 requests of 8,192 tokens each, so one GPU with 48 GB or more.

Meta's Llama 3 paper gives the 8B model 32 layers, a model dimension of 4,096 with 32 attention heads (a head size of 128) and 8 key/value heads. From that architecture the model has about 8.03 billion parameters.

Step Working Result
Weights at BF16 8.03 billion × 2 bytes 16.1 GB
Overhead 20% of the weights 3.2 GB
KV cache per token 2 × 32 × 8 × 128 × 2 bytes 131,072 bytes (0.13 MB)
Tokens in flight 16 requests × 8,192 tokens 131,072 tokens
KV cache 131,072 × 131,072 bytes 17.2 GB
Total 16.1 + 3.2 + 17.2 36.5 GB
GPU memory needed 36.5 ÷ 0.92 39.6 GB

One 48 GB GPU fits it, with room for about 23 requests at full length. One 80 GB GPU holds about 50. A 24 GB GPU does not, unless you quantize the weights or serve fewer, shorter requests.

How much memory does Llama 3.1 70B need?

About 212 GB at 16-bit for the same load, which in practice means four 80 GB GPUs, or two with 8-bit weights.

The 70B model has 80 layers, a model dimension of 8,192 with 64 attention heads (head size 128) and 8 key/value heads, which works out to about 70.55 billion parameters. Hugging Face rounds this to 71B.

Step BF16 weights FP8 weights
Weights 70.55 billion × 2 bytes = 141.1 GB 70.55 billion × 1 byte = 70.6 GB
Overhead (20% of weights) 28.2 GB 14.1 GB
KV cache per token 2 × 80 × 8 × 128 × 2 bytes = 327,680 bytes same
KV cache, 131,072 tokens in flight 42.9 GB 42.9 GB
Total 212.3 GB 127.6 GB
GPU memory needed (÷ 0.92) 230.7 GB 138.7 GB
80 GB GPUs 4 (3 would hold it, but see below) 2

Why four and not three? A model too large for one GPU is usually split with tensor parallelism, which divides each layer's attention heads across the GPUs. Sixty-four heads and eight KV heads divide evenly across 2, 4 or 8 GPUs, not 3. The vLLM documentation suggests pipeline parallelism, which splits by layers, when the GPU count doesn't divide the model evenly.

How many GPUs do you need for throughput?

As many copies of the model as it takes to serve your peak traffic within your latency target, which you can only learn by measuring.

Memory tells you the smallest unit that can run the model: one 48 GB GPU for the 8B example, a group of four 80 GB GPUs for the 70B at 16-bit. Then load test one unit with prompts like yours and read its total tokens per second. The cost calculator turns that figure into a count: monthly tokens ÷ (tokens per second × target utilization × seconds in a month), rounded up, plus spares. If one unit spans several GPUs, enter the group's price and throughput and read "GPUs" as groups.

How can you fit a model on fewer GPUs?

Shrink one of the three terms, and check the quality and latency cost of each change.

  1. Quantize the weights. FP8 halves weight memory and 4-bit quarters it, at some risk to quality. The post on quantization covers the methods.
  2. Quantize the KV cache. vLLM's --kv-cache-dtype accepts fp8, which halves the KV term.
  3. Cap the context. --max-model-len limits tokens per request. If your requests never approach the model's full 128K, a lower cap frees memory for more of them.
  4. Cap concurrency. Fewer requests in flight means less KV cache, at the cost of queueing at peaks.
  5. Use an engine that pages the KV cache. vLLM's PagedAttention (Kwon et al., 2023) reports "near-zero waste" in KV cache memory, so real requests shorter than the maximum leave room for more of them.

The formula above is a planning estimate for the worst case, every request at full length. Confirm it with a load test before you buy.

Hyperlake runs open models through KServe on Kubernetes, on compute you choose, and adds no compute markup; the sizing above applies whatever you use to serve the model.

Frequently asked questions

Can a 70B model run on a single GPU?

Only with heavy quantization. At 4-bit, Llama 3.1 70B's weights take about 35 GB, so with overhead an 80 GB GPU has roughly 31 GB left for KV cache: about 11 requests at 8,192 tokens. Check answer quality at 4-bit on your own prompts first.

Which matters more, model size or context length?

Both, because KV cache grows with every token in flight. One Llama 3.1 70B request at its full 128K context needs about 43 GB of KV cache at 16-bit, the same as 16 requests at 8K. For the 8B model, a single 128K request needs more memory for KV cache (17.2 GB) than for the weights (16.1 GB).

Do mixture-of-experts models follow the same formula?

For weights, use the total parameter count, not the active count, because all the experts are normally loaded into GPU memory even though only some run for each token. The KV cache term depends on the model's attention layout, so read its config for layers, KV heads and head size.

Why does the serving engine use almost all GPU memory when idle?

Engines such as vLLM reserve memory up front and fill what is left after the weights with KV cache blocks. vLLM's default share is 0.92 of GPU memory per instance, set with --gpu-memory-utilization.

Sources

Sources for the facts on this page, last checked October 9, 2026.

  1. Mastering LLM Techniques: Inference Optimization (NVIDIA Technical Blog, 2023) checked October 9, 2026
  2. The Llama 3 Herd of Models (Meta, arXiv 2407.21783), table of model hyperparameters checked October 9, 2026
  3. Transformer Math 101 (EleutherAI, 2023) checked October 9, 2026
  4. Engine arguments, including --gpu-memory-utilization and --kv-cache-dtype (vLLM documentation) checked October 9, 2026
  5. Parallelism and scaling (vLLM documentation) checked October 9, 2026
  6. Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., 2023) checked October 9, 2026
  7. Llama-3.1-70B-Instruct model card (Hugging Face) checked October 9, 2026

Start with a workload. Build the environment around it.

Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.