# Quantization Explained for LLM Inference

> Quantization explained: learn how lower-precision weights reduce LLM memory and bandwidth demands, and when to choose 8-bit, 4-bit, FP8, PTQ, or QAT.

Source: https://hyperlake.cloud/blog/quantization-explained
Published 2026-10-07 · by Hyperlake Team · Hyperlake

Video: [Watch: Quantization Explained (2:59)](https://www.youtube.com/watch?v=HVJVlNfeVg0)

Quantization reduces the numerical precision used to represent a model’s weights, commonly converting 16-bit floating-point values into 8-bit or 4-bit formats. The model retains the same basic structure and parameter count, but its weights occupy less memory and require less bandwidth to move, which can make inference faster and less resource-intensive.

Quantization matters because loading billions of weights from GPU memory can become the limiting factor in language model inference. Lower precision may let teams serve models with fewer accelerators or different classes of hardware, although the outcome depends on model size, workload, runtime, and hardware support. The video above walks through the core ideas.

## What is model quantization?

Model quantization represents weights, activations, or both with fewer bits than the original model used. A model trained or stored with 16-bit floating-point values may be converted to 8-bit integers, 4-bit integers, FP8, or another compact format.

The process maps high-precision values into a smaller set of representable values. Scales, zero points, and grouping methods help approximate the original numbers during computation. This mapping introduces rounding error, so quantization trades some numerical precision for lower memory use and potentially faster execution.

Quantization does not remove layers or parameters. In an idealized weight-only comparison, 4-bit weights require roughly one quarter of the raw storage used by 16-bit weights. Total runtime memory does not fall by exactly the same amount because activations, KV caches, temporary buffers, and serving software also consume memory.

This reduction can sometimes move a model from multiple GPUs to one GPU, or from a server accelerator to consumer hardware. Whether it fits still depends on parameter count, context length, batch size, concurrent requests, runtime overhead, and available memory.

## How do post-training quantization and quantization-aware training differ?

Post-training quantization converts an already trained model, while quantization-aware training exposes the model to reduced-precision behavior during training. PTQ is faster to apply and requires no additional training dataset, whereas QAT needs more compute but can preserve quality better at aggressive bit widths.

A typical PTQ process is:

1. Start with a trained higher-precision model.
1. Analyze or calibrate representative model values.
1. Convert values into the target low-precision format.
1. Evaluate quality and serving performance on the intended workload.

QAT simulates rounding and restricted numerical ranges during training or fine-tuning. The model can adjust its parameters to compensate for those effects. It is useful when a required low-bit deployment loses too much quality under PTQ, but the additional training makes it more expensive and operationally involved.

Deployment formats also differ in calibration, weight grouping, runtime support, and hardware compatibility. The [comparison of AWQ, GPTQ, and GGUF](https://hyperlake.cloud/blog/awq-gptq-and-gguf-quantization-methods-comparison) covers several practical approaches to post-training deployment.

![Diagram: PTQ and QAT compared by timing, compute requirements, and low-bit quality preservation.](https://hyperlake.cloud/blog/img/production/dbfe99eec70f2be1085bf2828d7376846e5420d7-1200x750.png?w=1600&fit=max&auto=format)

*PTQ prioritizes rapid conversion, while QAT adapts the model during training.*

## Should you use 8-bit, 4-bit, or FP8 quantization?

Choose a format by testing quality, memory use, latency, throughput, and compatibility on the intended production stack. Eight-bit formats are generally a conservative choice, 4-bit formats provide greater compression with more quality risk, and FP8 offers floating-point behavior on compatible accelerators.

Eight-bit quantization is often close to the higher-precision model on common inference tasks and is broadly supported by production hardware and frameworks. It reduces weight memory without restricting values as aggressively as 4-bit integer formats.

Four-bit quantization provides a larger reduction in raw weight storage and is supported by modern serving frameworks. It requires more careful calibration and evaluation because some models, tasks, or layers are more sensitive to reduced precision. Implementations may keep selected operations at higher precision to protect output quality.

Below 4 bits, degradation generally becomes harder to control, and the techniques needed to retain acceptable quality become more complex. The smallest format is therefore not automatically the most useful or economical production option.

FP8 retains a floating-point exponent and mantissa rather than using an integer representation. It is increasingly used for training and inference on compatible modern accelerators because it can reduce memory and bandwidth requirements while retaining numerical behavior suitable for many training workloads. Support still depends on the hardware, framework, kernels, and model architecture.

![Diagram: Four checks for choosing between 8-bit, 4-bit, FP8, and lower-precision model formats.](https://hyperlake.cloud/blog/img/production/6904e8c430124f57ac2fbee31e402e25a31b6e5d-1200x750.png?w=1600&fit=max&auto=format)

*Balance model quality, memory use, serving performance, and stack compatibility.*

## Why can quantization make inference faster?

Quantization can accelerate inference because hardware moves fewer bytes for each model weight. This is most helpful when memory bandwidth, rather than arithmetic capacity, is the binding constraint.

The model still performs essentially the same mathematical workload: quantization is not pruning, and it does not inherently reduce the parameter count. Its performance advantage mainly comes from storing and transferring compact values, with optimized kernels performing low-precision operations or reconstructing values efficiently.

Results depend on the complete serving stack. Hardware must support the selected data type, and the runtime needs suitable kernels that avoid excessive conversion overhead. Batch size, context length, KV cache pressure, and model architecture can also shift the bottleneck away from weight movement.

Teams should benchmark the exact model, format, accelerator, serving framework, and prompts planned for production. A broader [LLM testing and benchmarking process](https://hyperlake.cloud/blog/llm-testing-and-benchmarking-in-enterprise) should assess output quality alongside memory, latency, and throughput.

## Key takeaways

- Quantization lowers the bit width of model values without inherently changing the model’s parameter count.
- Four-bit weights use roughly one quarter of the raw storage of 16-bit weights, although total runtime memory includes other components.
- PTQ is faster to apply, while QAT can better preserve quality at aggressive precision levels.
- Eight-bit, 4-bit, and FP8 formats have different quality, portability, and hardware requirements.
- Quantization helps most when memory capacity or bandwidth limits inference performance.

## How Hyperlake helps

Hyperlake lets teams assemble and operate private AI environments with open or custom models, selected compute, evaluation, observability, and governed access to enterprise context. Open models can be served through KServe, subject to the deployment and validated integration, while workload-aware CPU and GPU capacity provides direct visibility into infrastructure usage and cost. To discuss model serving in your infrastructure or a client’s, [talk to our team](https://hyperlake.cloud/contact).

## Frequently asked questions

### Does quantizing an LLM always reduce its accuracy?

Quantization introduces approximation error, but that error does not always cause a meaningful decline on the target task. Eight-bit quantization often remains close to the original model, while 4-bit and lower formats require more careful validation. Teams should test representative prompts, domain data, output criteria, and the exact production runtime.

### Can a 4-bit model always run on one GPU?

No. Four-bit weights reduce raw model storage, but inference also needs memory for activations, KV caches, temporary buffers, framework overhead, and concurrent requests. Whether a model fits on one GPU depends on its parameter count, context length, batch size, quantization method, runtime implementation, and the accelerator’s available memory.

### Should a production team choose PTQ or QAT?

PTQ is usually the practical starting point because it converts an existing model quickly without another full training cycle. QAT is more appropriate when the required low-bit format causes unacceptable degradation under PTQ and additional training compute is available. Either choice should be validated against workload-specific quality and performance requirements.
