
KV cache routing directs an inference request to the server most likely to hold reusable key-value tensors for its prompt prefix. By preserving cache locality across replicas, it avoids repeating prefill computation for long, shared contexts. When no useful cached prefix exists, the router can fall back to conventional load-based placement.
This matters because enabling prefix caching on individual servers does not preserve its benefits across a distributed inference cluster. Routing strategy can therefore change the effective cost of otherwise identical deployments. The video above walks through the core ideas.
What is KV cache routing?
KV cache routing is a cluster-level technique that places requests according to the cache contents of inference replicas. It extends prefix caching beyond one server by making the routing layer aware of where reusable prompt state probably resides.
During transformer inference, the prefill phase processes input tokens and computes attention key and value tensors. If later requests begin with the same tokens, a server can reuse those tensors instead of computing the shared prefix again. The potential savings grow with the length and reuse frequency of that prefix.
Each inference replica normally maintains an independent, GPU-resident KV cache. A prefix cached by one server provides no benefit when the next matching request reaches another server that has never processed it. KV cache routing addresses this limitation through request placement; it does not create one shared cache across every GPU.
How does cache-aware routing work across replicas?
A cache-aware router identifies an incoming request’s prefix and determines which server is most likely to retain reusable state for it. It prefers that server when practical, then falls back to load-based placement when no useful cache state exists.
A typical sequence is:
- Identify the prefix. Derive a stable representation, often a hash, from the reusable part of the request.
- Check likely placement. Consult cache-state information that associates prefixes with inference replicas.
- Prefer a probable hit. Route the request to a replica that has likely retained the relevant KV tensors.
- Apply a fallback. If no useful state exists, consider queue depth, request count, capacity, or another load signal.
Production routing must balance locality against congestion. Sending every related request to one replica could create a queue while other replicas remain available, so cache affinity should account for current load. It also complements techniques such as continuous batching, which schedules concurrent work after requests reach a server.

Why does ordinary load balancing reduce cache hits?
An ordinary load balancer considers server load but not cached prefixes. By spreading matching prompts among replicas, it can force several servers to perform the same prefill computation and dilute the locality that prefix caching requires.
For example, a cluster may receive many requests containing the same long system prompt. A cache-blind policy can send consecutive requests to different replicas because each appears less busy. Every selected server then computes and stores the shared prefix independently.
A cache-aware router instead concentrates similar requests on servers that have already completed the prefill work. It uses load-based routing when no useful cache exists or when the preferred server is unsuitable. Consequently, deployments with the same model, inference framework, and GPU capacity can have different effective costs because one repeats more prefill work. Serving architecture is therefore part of the economics of large language models.

Which workloads benefit from KV cache routing?
KV cache routing helps most when requests repeatedly contain long, stable prefixes. It offers no meaningful reuse when every prompt is unique, regardless of how sophisticated the routing infrastructure is.
Likely candidates include workloads with:
- Long system prompts shared across requests.
- Repeated document contexts used for several questions or processing steps.
- Common tool descriptions included in agent requests.
- Stable instruction blocks shared across users, sessions, or jobs.
Prefix-cache hit rate is workload-specific rather than one universal cluster metric. A single service may combine highly reusable agent prompts, changing documents, and unique user requests, so an aggregate measurement can conceal important differences.
Teams should measure hit rate by workload type, along with shared-prefix length, reuse frequency, and request distribution across replicas. They can then compare avoided prefill work with the operational complexity of maintaining cache-state-aware routing. This determines whether the optimization is justified for a particular deployment rather than assumed useful everywhere.
Key takeaways
- Prefix caching reuses key-value tensors for repeated prompt prefixes instead of recomputing them for every request.
- Independent inference replicas do not automatically share their GPU-resident KV caches.
- Cache-aware routing prefers the server most likely to hold a requested prefix.
- Load-based fallback remains important when no useful state exists or the preferred replica is congested.
- Teams should measure prefix-cache hit rate separately for each workload type.
How Hyperlake helps
Hyperlake lets teams assemble, deploy, govern, and operate model services and supporting applications on Kubernetes-based infrastructure they or their clients control. Its modular architecture can run open models through KServe, subject to the deployment and validated integration, while shared controls support observability and lifecycle operations. To discuss cache-aware routing within a sovereign inference environment, talk to our team.
Frequently asked questions
Does KV cache routing share GPU memory between inference servers?
No. Each inference server generally retains its own GPU-resident KV cache, and another replica cannot automatically use that state. KV cache routing improves reuse by sending a matching request to the replica most likely to have already processed and retained the required prefix.
Can KV cache routing help when every prompt is unique?
KV cache routing offers little benefit when requests do not share reusable prefixes. The router cannot produce a cache hit if no replica has previously processed the same prefix. In that situation, load, queue depth, and available capacity are more useful placement signals because every server must perform the prefill computation.
What should teams measure before adopting cache-aware routing?
Teams should measure prefix-cache hit rate by workload type, shared-prefix length, reuse frequency, prefill work, and request distribution across replicas. These measurements reveal whether repeated system prompts, document contexts, or tool definitions create useful locality. A cluster-wide average can hide major differences among workloads.
How is prefix caching different from response caching?
Prefix caching stores intermediate attention key-value tensors so the model can avoid repeated prefill computation while still generating a new response. Response caching stores and returns an already completed output for a matching request. KV cache routing concerns the placement of active inference work, not retrieval of a finished answer.


