Two requests can contain almost the same 20,000-token prompt and still reach the first generated token at very different times. On the first request, the inference server may have to run the entire prompt through the model. On the next, much of that work may already exist in GPU memory.
The difference is prefix caching. Instead of treating every incoming prompt as completely new, a capable inference engine can recognize an identical beginning, find the KV-cache state produced when that prefix was processed earlier, and continue from there.
This does not make long prompts free. It turns repeated prompt computation into a memory-management problem: the reusable state has to exist, match, remain resident and survive competition from other requests.
- Prefix caching reuses previously computed KV state for an identical prompt prefix instead of performing the same prefill work again.
- The main direct benefit is reducing repeated prompt processing, so time to first token can improve while decode speed remains largely a separate issue.
- A cache hit requires compatible prefix state; a prompt that merely looks similar to a person is not necessarily reusable to the runtime.
- Cached state occupies finite memory, so reuse, allocation and eviction become part of inference-server capacity planning.
- In multi-tenant systems, shared prefix caching also creates isolation and security considerations rather than being a purely performance-oriented feature.
- The expensive prompt may already have been processed
- What the server is actually reusing
- “Almost the same prompt” is not the same thing as a cache hit
- A cache hit mostly attacks the wait before generation
- The speedup has to live somewhere
- RAG and agents can create unusually valuable prefixes
- Concurrency turns prefix caching into a placement problem
- Shared caches also create a security boundary
- The fastest prompt is sometimes the one the GPU does not process twice
The expensive prompt may already have been processed
Before an LLM can generate its first output token, it has to process the input that came before it. A system prompt, chat history, retrieved documents, tool definitions and the user’s latest message can all become part of that input.
This is the prefill phase. As explained in our guide to prompt processing versus token generation, it is a different workload from decoding new output tokens one at a time.
Now consider an application that repeatedly sends the same large system prompt followed by a different user question. Without reusable prefix state, the server processes that common beginning again for each request.
Prefix caching asks a simple question before doing that work: has this exact beginning already been processed in a form the engine can reuse?
If the answer is yes, the server can recover the corresponding KV-cache state and perform new computation only where the reusable prefix ends.
What the server is actually reusing
Prefix caching is easy to misunderstand as a text cache. The important reusable object is not simply a stored copy of the words in the prompt.
During transformer inference, attention layers produce key and value states for previously processed tokens. A KV cache keeps those states so the model does not have to reconstruct all previous attention information from scratch for each subsequent generation step.
That same underlying state can become useful across requests. If another request begins with a compatible prefix, the serving engine can reuse the already-computed KV state associated with that prefix rather than rerunning the entire shared portion through the model.
The distinction matters physically. Compute that would have taken place on the accelerator is replaced by a lookup and reuse operation against memory already occupied by cached state.
That also explains why prefix caching is tied so closely to memory management. The reusable result has to live somewhere.
Prefix caching converts repeated GPU computation into a state-reuse problem. The server spends memory holding KV blocks so a later request can avoid part of its prefill computation. The useful resource is therefore not just GPU compute or VRAM in isolation, but the relationship between repeated prefixes, available cache capacity and the engine’s ability to find and retain reusable state.
“Almost the same prompt” is not the same thing as a cache hit
Humans can look at two prompts and decide that they are basically identical. A cache manager needs a stricter definition.
vLLM’s current prefix-caching design manages KV state in blocks. Full blocks can be identified using hashes derived from their token content, the preceding prefix and additional information required to distinguish otherwise incompatible states.
The consequence is important: reuse depends on the actual processed prefix, not on semantic similarity.
Move changing information toward the beginning of a request and the common prefix may end earlier. Keep stable instructions and shared material before request-specific content and there may be more reusable state, assuming the application’s template and serving engine preserve a compatible prefix.
This is why prompt construction can have a physical performance consequence. Reordering logically equivalent application content can change where two requests stop sharing reusable computation.
| Request pattern | Reusable prefix potential | Physical consequence |
|---|---|---|
| Large stable system prompt followed by changing user input | Potentially high | Previously computed prefix state may replace repeated prefill work. |
| Same long document followed by different questions | Potentially high if the rendered prefix remains identical | The document portion may not need to be processed from the beginning for every request. |
| Changing metadata inserted near the start | Potentially lower | The shared prefix can terminate earlier, leaving more tokens to process again. |
| Completely unrelated prompts | Low | Cached blocks provide little reusable prefill work for the new request. |
A cache hit mostly attacks the wait before generation
Prefix caching does not make every phase of inference faster.
Its direct target is redundant processing of input that the model has already seen in an identical reusable prefix. Avoiding that work can reduce the prefill portion of latency and therefore improve time to first token.
Once the server reaches new input and begins generating new output, it still has work to perform. The target model still has to process uncached tokens and decode the response.
This is why a workload can show a large difference in initial responsiveness without showing the same proportional change in the rate at which the answer streams afterward.
The distinction is especially important when benchmarking. A single “tokens per second” number can hide the effect entirely because prefix caching primarily changes how much input computation has to happen before generation.
The speedup has to live somewhere
The phrase “cache hit” can make the optimization sound almost free. Physically, the server has exchanged recomputation for retained state.
KV cache already matters during ordinary generation because previous attention state consumes memory as sequences grow. Our article on context windows and VRAM explains why that state becomes part of the inference memory budget.
Prefix reuse adds another dimension: completed or otherwise reusable blocks can remain available so future requests can reference them.
But accelerator memory is finite. A serving engine cannot preserve every historical prefix indefinitely while also allocating memory for model weights, active sequences and runtime requirements.
Eventually blocks have to become available for new work. In vLLM’s block-based design, the cache manager maintains a pool of blocks and a free-block queue. Cached blocks that are no longer actively referenced can remain available for reuse, but they can also be reclaimed when new allocations require capacity.
The performance of prefix caching therefore depends partly on locality: whether useful prefixes are requested again while their state is still available.
A large cache does not guarantee a high cache-hit rate. If requests rarely share prefixes, cached state may consume capacity without eliminating much repeated work. Conversely, a workload with a stable long prefix can make reuse disproportionately valuable because the same expensive prefill region appears again and again.
RAG and agents can create unusually valuable prefixes
Some workloads naturally repeat large sections of input.
An internal assistant may attach the same policy document to many questions. An agent may send a long stable system prompt and the same tool definitions on every turn. A document-analysis service may repeatedly query one large source while changing only the question near the end.
These patterns are interesting because the repeated text appears before the changing part of the request.
If that common region remains token-for-token compatible with cached state, the server has an opportunity to reuse substantial previous computation. If dynamic information is injected early in the prompt, the reusable region can become much shorter.
This means application architecture can affect inference infrastructure without changing the model. Two applications can send roughly the same number of tokens to the same model on the same GPU yet create different amounts of reusable work simply because their prompts are assembled differently.
| Workload | Repeated region | What prefix caching can change | What it does not remove |
|---|---|---|---|
| Long system instructions | Stable instructions at the beginning | Repeated prefill for the matching prefix | New user input and output generation |
| Document Q&A | Same document before different questions | Repeated processing of the reusable document prefix | Changed suffix and decode work |
| Agent with stable tools | Instructions and tool definitions | Repeated computation for the stable beginning | Dynamic tool results and subsequent generation |
| Unrelated chat requests | Little beyond a small common template | Potentially little | Most request-specific inference work |
Concurrency turns prefix caching into a placement problem
A single inference process can decide whether it already holds a useful prefix. A fleet of inference servers has a harder problem: the reusable state may exist, but on a different machine.
Suppose one replica has already processed a long shared document. Sending the next related request to that replica may allow reuse. Sending it to another replica can force the same prefix through prefill again because the relevant cached state is not local there.
At that scale, routing and cache locality become connected.
This extends the capacity problem described in How Many Users Can One GPU Actually Serve? Request count alone does not determine how much GPU work arrives. The amount of reusable state can change the actual computation associated with apparently similar traffic.
A routing policy optimized only for distributing request counts evenly can therefore miss another resource: already-computed state.
Measure cache behavior together with latency. Compare first-request and repeated-prefix TTFT, record how much input is actually reused, and watch what happens as concurrency and cache pressure increase. If TTFT improves only while a prefix is warm and then regresses under load, the useful question is no longer whether caching works. It is whether reusable state survives long enough to matter for the real traffic pattern.
Shared caches also create a security boundary
Performance is not the only reason a production system needs to know who can reuse which prefix.
A cache hit changes execution behavior. In a shared service, differences in response timing can potentially reveal information about whether another request previously created matching cached state.
Current vLLM documentation addresses this with optional cache salting. A salt becomes part of the cache identity so that only requests using the same value can reuse those blocks. This creates a way to separate cache-sharing groups rather than allowing every tenant to participate in one unrestricted prefix namespace.
That introduces a real trade-off. Broader sharing can create more reuse opportunities; stronger isolation deliberately prevents some of those opportunities.
Prefix caching is therefore not simply an “on equals faster” switch. In a multi-user deployment, its scope belongs in the system’s security and tenancy design as well as its performance configuration.
A useful first-party test would use one fixed model and GPU with a long stable prefix followed by a short changing suffix. Measure the first request, repeat the workload with the prefix warm, then increase unrelated traffic until cache pressure changes the result. TTFT, reused-prefix length and cache behavior would reveal when stored state actually replaces GPU prefill work.
The fastest prompt is sometimes the one the GPU does not process twice
A long input normally looks like work waiting to happen. Prefix caching changes that assumption.
Once an inference engine has computed reusable state for the beginning of a request, the same tokens can represent something physically different the next time they arrive. Instead of another full stretch of GPU prefill, they can become a lookup into state already held by the serving system.
That advantage survives only while the prefix still matches, the cache state remains available and the request reaches hardware that can reuse it.
For repeated long prompts, then, counting input tokens is not enough. The more useful question is how many of those tokens still require the GPU to do the work again.







