Four people are talking to the same local LLM. One asks a short question and gets an answer in seconds. Another requests a long explanation. A third pastes several pages of text. The fourth request arrives while all of that work is still running.
A single-user benchmark tells you surprisingly little about what happens next. The GPU does not simply finish one conversation and move to the next, and a modern inference server does not have to freeze a fixed group of requests until the slowest one is done.
Instead, a scheduler can repeatedly rebuild the work sent to the accelerator. Finished sequences leave. Waiting requests enter. Long prompts compete with ongoing token generation. KV-cache blocks are allocated and released as conversations grow and finish.
That mechanism is continuous batching. It is one of the reasons an LLM server can turn the same physical GPU into a much more useful shared resource.
- Continuous batching changes the active batch while generation is running instead of waiting for a fixed batch to finish.
- The scheduler is balancing GPU work, request latency and a finite KV-cache memory budget at the same time.
- Higher aggregate throughput does not guarantee lower latency for every request.
- Prompt length matters because prefill work can be much larger than a single decode step.
- The useful measurements are TTFT, inter-token latency, queue time, throughput and KV-cache pressure — not GPU utilization alone.
- Why a fixed batch is awkward for text generation
- Prefill and decode compete for the same machine
- The batch is dynamic, but VRAM is still finite
- Why PagedAttention changed LLM serving
- A busy GPU is not enough evidence
- Throughput and responsiveness can move in different directions
- The scheduler is part of the machine
Why a fixed batch is awkward for text generation
Batching is attractive because GPUs are built to execute large amounts of parallel numerical work. Giving the accelerator more useful work at once can improve utilization and aggregate throughput.
Text generation creates a complication: requests rarely have identical lifetimes.
One user may generate 40 tokens. Another may generate 400. A third may stop early after producing an end-of-sequence token. If those requests are treated as a fixed batch that must remain together until every sequence finishes, completed requests can leave unused capacity behind while the longest request continues.
Continuous batching removes that requirement. Hugging Face’s current Transformers documentation describes the batch as being dynamically rescheduled during generation: when one request finishes, another can enter rather than waiting for the original group to complete.
This is not just queue management around the GPU. The scheduler is deciding what work becomes the GPU’s next workload.
That distinction becomes important when the server mixes two very different phases of inference: prompt processing and token generation. Our guide to prompt processing versus token generation explains why these phases have different performance characteristics.
Prefill and decode compete for the same machine
When a new prompt arrives, the model first has to process the input tokens. This is the prefill phase. Once the first output token is ready, generation moves into decode, where the model repeatedly produces subsequent tokens while reusing state stored in the KV cache.
A decode iteration may involve only a small amount of new token work for each active sequence. A large incoming prompt can represent thousands of tokens of prefill work.
Put both on one server and the scheduler has a real trade-off. Admit large prefills aggressively and the machine can process incoming prompts efficiently, but existing streaming requests may experience less predictable token delivery. Protect decode work too strongly and new requests can spend longer waiting for their first token.
This is why serving performance cannot be reduced to a single tokens-per-second number.
| Work entering the server | What the machine must do | Main resource pressure | What the user notices |
|---|---|---|---|
| Short new prompt | Run a relatively small prefill, then begin decode | Compute + new KV state | Time to first token |
| Long new prompt | Process many input tokens before generation starts | Prefill compute + KV growth | Potentially longer TTFT |
| Existing streaming request | Advance decode and append new KV state | Memory bandwidth, compute and cache capacity | Inter-token latency |
| Many simultaneous conversations | Schedule multiple active sequences repeatedly | KV capacity + scheduler admission + GPU work | Queueing and tail latency |
A software request eventually becomes allocations and tensor operations on a physical accelerator. Model weights occupy GPU memory. Active sequences consume KV-cache capacity. Input tensors and runtime buffers need space. The scheduler cannot admit infinite work simply because HTTP requests are cheap to create.
The batch is dynamic, but VRAM is still finite
Continuous batching solves a utilization problem. It does not solve the memory-capacity problem.
Each active sequence needs state associated with its context. As our article on context windows and VRAM explains, KV-cache requirements depend on the sequence length, model architecture, cache representation and other runtime details.
Serving engines therefore need a way to allocate cache memory as requests enter, grow and finish. Modern systems commonly divide the KV cache into blocks rather than reserving one rigid contiguous region for every possible sequence.
NVIDIA’s TensorRT-LLM documentation describes its KV cache as pools of blocks assigned to requests as needed. Hugging Face’s continuous-batching implementation likewise exposes a paged cache and explicitly treats the token-batch budget and number of available cache blocks as competing users of GPU memory.
This turns VRAM into something closer to a managed capacity pool.
Paged KV-cache management can reduce memory waste and make dynamic request membership practical. It does not make context free. Enough simultaneous long sequences can still fill the available cache pool, forcing the scheduler to wait, reject, evict or preempt work depending on the serving engine and configuration.
Why PagedAttention changed LLM serving
The idea behind paged attention is closely related to virtual-memory paging in operating systems: logical sequence state does not have to correspond to one large contiguous physical allocation.
The original vLLM PagedAttention work focused on a major serving problem: KV-cache memory grows dynamically and can be wasted through fragmentation and redundant allocation. Breaking that state into manageable blocks allows the serving system to use GPU memory more flexibly.
That matters to continuous batching because a changing batch also means a changing set of active contexts. Requests arrive at different times, have different prompt lengths, generate different numbers of tokens and release their cache state at different moments.
The scheduler and cache manager therefore work together. One decides which requests should advance. The other determines whether the physical memory required for those active sequences is available.
A busy GPU is not enough evidence
Suppose continuous batching raises GPU utilization from an obviously underused state to something that looks much healthier. That is useful, but it does not tell you whether the service is good.
The machine and the user observe different things.
vLLM exposes these dimensions separately. Its current production metrics include running and waiting request counts, KV-cache usage, TTFT, inter-token latency, queue time, prompt-token counts and generation-token counts.
That list is revealing. A real serving engine does not consider “tokens per second” sufficient to describe system health.
This also explains why the answer to how many users one GPU can serve depends on the workload. Twenty mostly idle chat sessions are not the same physical load as twenty simultaneous long-context generations.
Watch request arrival rate, queue depth, TTFT, inter-token latency and KV-cache occupancy together. Then correlate them with prompt and output lengths. If queue time rises while KV-cache utilization approaches its limit, the problem is different from a server where cache capacity is comfortable but long prefills dominate TTFT.
Throughput and responsiveness can move in different directions
A scheduler can make a GPU complete more total work by keeping it fed with a larger and more varied set of requests. That is valuable when the objective is aggregate throughput.
But interactive AI has another objective: users want responses to start quickly and continue smoothly.
Those goals can conflict. Admitting more work can improve utilization while increasing contention. Larger prefill batches can process incoming prompts efficiently while consuming memory and compute that could otherwise advance existing decode sequences. Reserving capacity for active decodes can protect responsiveness while making waiting requests spend longer in the queue.
The right configuration therefore comes from the workload rather than from a universal “best batch size.”
| If this gets worse | Inspect first | Possible mechanism |
|---|---|---|
| TTFT under load | Queue time, prompt lengths, prefill time | Incoming work is waiting or large prefills dominate scheduling |
| Streaming smoothness | Inter-token latency and active request count | Decode sequences are competing for iterations |
| Concurrency ceiling | KV-cache usage and context lengths | Active sequence state is exhausting the cache budget |
| GPU utilization | Running/waiting requests and batch-token behavior | The scheduler may not have enough eligible work to keep the accelerator occupied |
A useful first-party test would hold the model, GPU and serving engine constant while increasing request concurrency and changing the mix of short and long prompts. The goal would not be to find one universal “users per GPU” number, but to show where throughput gains begin to trade against interactive latency.
The scheduler is part of the machine
It is easy to picture an AI server as a model sitting on a GPU and answering requests. Under load, that picture is incomplete.
Between the request queue and the accelerator sits a scheduling system deciding which tokens get compute time and which contexts get memory. Continuous batching makes that decision repeatedly so completed work can leave and new work can enter without waiting for an artificial fixed batch boundary.
That can make one GPU serve shared workloads far more effectively. But it also makes performance a systems problem. Prompt length, decode activity, queueing and KV-cache pressure interact inside the same physical memory and compute budget.
If a shared LLM server feels slow, the useful question is no longer simply whether the GPU is fast enough. Look at what the scheduler is asking that GPU to do — and what it had to leave waiting.
Technical basis: current Hugging Face Transformers continuous-batching documentation, vLLM scheduler and production-metrics documentation, NVIDIA TensorRT-LLM KV-cache documentation, and the original vLLM PagedAttention paper.







