KV Cache Quantization Explained: How FP8 Fits More LLM Context Into VRAM

Servers

A shared LLM server is approaching its limit. The model weights still fit. The GPU has not reported an out-of-memory error. But long conversations and concurrent requests have filled most of the remaining VRAM with KV-cache state. Requests begin waiting, and the serving engine moves closer to the preemptions and recomputation described in our investigation of what happens when the KV cache fills up.

One configuration change appears to offer an escape: store that cache in FP8 instead of BF16 or FP16. The nominal bytes per cached value fall from two to one. More tokens can remain resident on the same GPU.

That sounds like free capacity. It is not. KV-cache quantization changes how the server represents its working memory, how attention kernels read it and, in some implementations, the precision used during attention itself. The useful question is not whether FP8 saves memory. It does. The question is what the server gives up to obtain that space—and whether the real workload benefits.

Key Takeaways
  • KV-cache quantization is separate from model-weight quantization. A quantized model can still use a BF16 cache, and a BF16 model can use an FP8 cache.
  • Moving from a two-byte cache format to a one-byte format can approximately halve the raw KV payload, but scales, block allocation and runtime overhead remain.
  • The saved VRAM can support more active tokens, longer contexts or greater concurrency. It does not guarantee twice the throughput.
  • Kernel support, attention architecture and calibration determine whether the memory win arrives with acceptable latency and quality.
  • Measure cache pressure, preemptions, tail latency and task quality together. GPU utilization alone cannot validate the change.

The cache is not the model

Weight quantization and KV-cache quantization act on different objects. Model weights are the learned parameters loaded for inference. They are largely static while the server is running. The KV cache is generated from the prompts and tokens that are currently being processed. It grows and shrinks with the workload.

This distinction matters operationally. A server may run four-bit model weights and still maintain its KV cache in BF16. Another deployment may keep BF16 weights but store the cache in FP8. The first choice determines much of the model’s baseline memory footprint; the second changes how much working state each active sequence consumes.

That is why a model can fit comfortably at startup and encounter memory pressure later. The weights have not changed. The number of cached tokens has.

KV bytes ≈ 2 × layers × KV heads × head dimension × cached tokens × bytes per value The leading 2 represents keys and values. Real runtimes also introduce block, alignment, scale and allocator overhead.

For an illustrative model with 32 layers, eight KV heads and a head dimension of 128, a BF16 cache uses 128 KiB of raw KV data per cached token. One 32,768-token sequence therefore represents about 4 GiB before runtime overhead. Storing the same values in a one-byte FP8 format reduces the raw payload to about 2 GiB.

This is a calculation, not a benchmark. Actual consumption depends on the model’s attention design, cache block size, padding, parallelism and implementation. Models using grouped-query or multi-query attention can require fewer KV heads than conventional multi-head attention, which changes the result before quantization is considered.

Cache format Nominal payload Primary advantage Main risk
BF16 or FP16 2 bytes per value Broad, mature kernel support and a high-fidelity baseline Large cache footprint under long contexts or concurrency
FP8 1 byte per value Roughly half the raw KV payload of a 16-bit cache Scale selection, backend compatibility and model-dependent quality changes
Four-bit cache 0.5 byte per value before metadata Greater theoretical memory reduction More aggressive information loss and less uniform production support
Two-bit research methods 0.25 byte per value before metadata Very high compression in validated research settings Specialized quantizers, kernels and model-specific validation requirements

What the runtime changes

A BF16 cache stores each key and value with a relatively wide numerical representation. FP8 has fewer bits, so the runtime must map a range of higher-precision values into a smaller set of representable values. A scale controls that mapping.

When new tokens enter the cache, their key and value tensors are quantized and written into VRAM. During later attention steps, the engine either converts those values for computation or uses a kernel capable of operating on the quantized representation. The exact path depends on the runtime and attention backend.

Attention state New key and value tensors are produced for the current token.
Scale and quantize Higher-precision values are mapped into the chosen low-precision range.
Store in VRAM The smaller representation allows more token state to remain resident.
Read for attention A compatible kernel consumes or converts the cached values during decoding.
Physical Layer

The capacity gain exists because fewer bytes cross the GPU’s memory system for each cached value. That can reduce VRAM occupancy and, on a suitable kernel path, the memory traffic required during decoding. But quantization also introduces scale metadata and conversion or low-precision arithmetic. The hardware still performs work; the balance of that work has changed.

The saved VRAM has three possible destinations

First, it can support a longer context for one request. Second, it can keep more requests active at the same time. Third, it can provide headroom that prevents the scheduler from reaching the point where active state must be preempted.

These outcomes are not interchangeable. A single long conversation, a batch of short requests and a mixed production workload create different allocation patterns. The same nominal memory reduction can improve maximum context in one test and concurrency in another.

Nor should “half the cache bytes” be translated directly into “twice the throughput.” Throughput may still be limited by model-weight traffic, compute, scheduling, prefill work or the cost of quantization. If cache pressure was not the active bottleneck, reducing it may produce little visible improvement.

Reality Check

KV-cache quantization is most valuable when the unquantized cache is large enough to constrain the workload. If the server handles short prompts at low concurrency and already has comfortable VRAM headroom, the smaller cache may solve a problem that the machine does not have.

The difficult part is choosing what “close enough” means

Quantization rounds values. A scale that covers a wide numerical range reduces saturation but leaves fewer representable steps for ordinary values. A narrow scale preserves more detail near the centre but may clip outliers. The best mapping can vary between models, layers and attention heads.

Current vLLM documentation distinguishes per-tensor scaling from per-attention-head scaling. Per-tensor quantization uses one scale for each Q, K or V tensor. Per-head scaling is more granular, but currently requires the supported FlashAttention and calibration pathway described by the project.

Calibration estimates those scales from representative data. That word matters. A calibration set full of short conversational prompts may not expose the numerical behaviour of a long-document workload. A model can pass a small perplexity check and still lose accuracy on retrieval, reasoning or another task that depends on precise long-context attention.

The KIVI research project reached an even lower two-bit cache by treating keys and values differently: keys were quantized per channel and values per token. Its published results are evidence that aggressive compression can work under a carefully designed method—not evidence that every generic two-bit switch is safe for every model.

Technical Note

In its evaluated configurations, the KIVI paper reported up to a four-times larger batch and higher serving throughput while preserving near-baseline quality for the tested Llama, Falcon and Mistral models. Those are first-party research measurements tied to KIVI’s quantizer, implementation, models and workloads. They should not be treated as a universal production multiplier.

“Supports FP8” does not describe one identical path

Current vLLM supports FP8 KV-cache formats and offers uncalibrated and calibrated scaling paths. With its FlashAttention 3 backend, the project documents a path where attention operations also run in the quantized domain rather than treating FP8 only as a storage format.

That distinction changes both performance and risk. In the vLLM team’s 2026 evaluation, validated FP8 paths reduced cache memory and improved several long-context, memory-bound workloads. The same report identified exceptions: short contexts where fixed overhead may dominate, 256-dimensional attention heads with prefill sensitivity, hybrid-attention models with small sliding-window layers, and models that showed persistent uncalibrated quality loss.

TensorRT-LLM separately exposes FP8 KV-cache configuration and documents NVFP4 cache support for compatible workflows. Hugging Face Transformers offers a quantized cache for memory-constrained generation but explicitly warns that it can harm latency when contexts are short and sufficient GPU memory is already available.

The safe conclusion is narrower than “FP8 is faster.” Runtime version, GPU architecture, model structure and selected attention kernel are part of the result.

Observed situation What it suggests Next action
Cache utilization approaches its limit and preemptions appear KV capacity is affecting serving behaviour Compare the native cache with a supported FP8 configuration
Contexts are short and cache utilization stays low KV memory is probably not the current bottleneck Keep the higher-precision baseline unless testing shows another benefit
FP8 improves concurrency but TTFT becomes worse The selected prefill or attention path may have additional overhead Inspect head dimension, backend and workload phase separately
Latency improves but long-context task scores fall The scales or quantization granularity may be unsuitable Calibrate on representative data or leave sensitive layers unquantized
GPU utilization is high but useful output barely changes The additional activity may not be advancing more user-visible work Measure useful output throughput, preemptions and tail latency together

Measure the capacity win and its price

A convincing test needs more than peak tokens per second. Compare the same model, prompts, output limits, request arrival pattern and runtime version. Change only the cache format and the calibration strategy.

Track the point where the system first queues or preempts work. Then examine user-facing latency and a task-level quality set. Our guide to LLM inference metrics explains why aggregate throughput can improve while individual requests become less predictable.

Cache pressure Used KV blocks, active tokens and the remaining capacity before admission changes.
Preemptions Whether requests are being interrupted or forced to repeat previously completed work.
TTFT and ITL Prompt-phase waiting and the spacing between generated tokens, including tail percentiles.
Task quality Long-context retrieval, reasoning and application-specific correctness on representative prompts.
What to Measure

Record p50 and p95 latency, not only averages. Count completed output tokens rather than all GPU activity. Compare quality at several context lengths, and correlate every latency change with cache utilization, running requests, waiting requests and preemptions.

When FP8 is a reasonable starting point

Test it when long contexts or concurrent sequences consume a material share of VRAM; when preemptions appear before compute is saturated; or when capacity planning shows that KV state, rather than weights, controls how many requests can remain active.

Keep a BF16 or FP16 baseline when the workload is small, short-context and latency-sensitive. Benchmark carefully when the model uses unusual attention geometry, hybrid sliding-window layers or a backend that has not been validated for the target GPU.

And do not use the configured maximum context as the workload by default. The practical context limit should still reflect what users need. Our guide to context windows and VRAM explains why advertising a large maximum and operating efficiently at that maximum are different claims.

il.digital Lab

A useful first-party test would isolate cache precision

Run one model on one fixed GPU with identical request traces, comparing its native KV-cache dtype with uncalibrated FP8 and a calibrated FP8 configuration. The test should continue beyond the first capacity warning so it captures queueing, preemption and tail-latency behaviour.

Question How much additional useful concurrency does FP8 create before service quality degrades?
Controlled Model, GPU, runtime, prompts, sampling, arrival pattern and output limits.
Variable Native cache, uncalibrated FP8 and calibrated FP8.
Measure Cache use, preemptions, TTFT, ITL, throughput, power and task quality.
Workloads Short chat, long-document retrieval and mixed concurrent requests.
Charts Concurrency against p95 ITL, useful throughput and quality score.

Treat cache precision as serving policy

KV-cache quantization is not merely a memory-saving checkbox. It decides how faithfully the server preserves the intermediate state of every active conversation, how much of that state fits beside the model and which kernel path the GPU must execute.

The right configuration is the one that keeps more useful work resident without hiding the cost in slower first tokens, unstable long-context behaviour or degraded answers. Measure the point where memory pressure begins, change the representation, and verify the whole service again. A smaller cache is valuable only when the server becomes better at serving users—not simply better at filling VRAM differently.

Rate article
Add a comment