A large language model can have a powerful GPU almost to itself and still produce a single conversation one token at a time. The accelerator finishes one decode step, the next token becomes known, and only then can the next step begin.
That dependency is awkward for hardware designed to perform enormous amounts of parallel work. During low-concurrency inference, the bottleneck can be moving a large model through the GPU for each small piece of new output rather than exhausting all of the accelerator’s arithmetic capability.
Speculative decoding attacks that mismatch in an unusual way. It deliberately performs extra work with a cheaper predictor so the expensive target model can validate several possible future tokens together. When enough predictions survive verification, the server advances through the response with fewer expensive sequential target-model steps.
The trick is not free. Its value depends on the draft, acceptance rate, workload and how busy the GPU already is.
- Ordinary autoregressive decoding has a sequential dependency: the next output token depends on the state produced by the previous one.
- Speculative decoding uses cheaper work to propose future tokens, then lets the target model verify several candidates together.
- Rejected speculative tokens represent work that did not advance the final output.
- The technique is especially attractive when decoding is memory-bound and the GPU still has unused compute capacity.
- At high concurrency, speculative work can compete with real requests for compute, so lower single-request latency does not automatically mean higher server throughput.
- The expensive part is doing it again and again
- A smaller model guesses ahead
- Why this can be lossless rather than approximate
- Speculation spends compute to save sequential time
- The same optimization can behave differently under load
- There is more than one way to speculate
- The fastest token may be the target-model step you avoid
The expensive part is doing it again and again
After a prompt has been processed, an autoregressive language model predicts the next token. That token becomes part of the sequence, and the model runs again to determine what follows it.
This is why output generation behaves differently from prompt processing. Existing input tokens are already known and can be processed with substantial parallelism. Future output tokens are not known yet.
Our guide to prompt processing versus token generation explains this split in detail. Speculative decoding specifically targets the sequential character of the decode side.
The interesting hardware detail is that a decode step does not necessarily use every capability of a modern accelerator efficiently. For low-batch inference, repeatedly accessing the target model’s weights and memory can dominate while significant arithmetic capacity remains available.
The optimization trades one physical resource for another. Instead of asking the large target model to perform as many sequential decode iterations, the system spends additional compute proposing and verifying candidate tokens. It works best when that extra arithmetic is cheaper than the target-model steps it eliminates.
A smaller model guesses ahead
The classic version uses two models. A smaller, faster draft model generates a short continuation. The larger target model then evaluates those proposed positions and determines which can be accepted.
If several candidates survive, generation has advanced by several tokens after one expensive target-model verification step instead of requiring the same number of ordinary sequential target-model steps.
If the draft is frequently wrong, much of that advantage disappears. The server still paid to generate the proposals, but rejected candidates do not become useful output.
This makes acceptance behavior central to the economics of speculation. A fast drafter that predicts poorly can create cheap but mostly useless work. A stronger drafter may produce better candidates but consume more time and resources before verification even begins.
| Stage | Physical work | Why it exists | What can go wrong |
|---|---|---|---|
| Draft | Cheaper model or prediction mechanism runs ahead | Create candidate future tokens inexpensively | Drafting itself becomes too expensive |
| Verification | Target model evaluates candidate positions | Preserve the target model’s output behavior | Verification cost dominates the saving |
| Acceptance | Accepted candidates extend the real sequence | Skip some sequential target-model decode steps | Low acceptance means little forward progress |
| Rejection | Unaccepted speculative work is discarded | Prevent incorrect draft predictions becoming output | Compute was spent without advancing useful generation |
Why this can be lossless rather than approximate
The word “guess” can make speculative decoding sound like a quality shortcut. The important idea is that the draft does not get final authority over the output.
In the original speculative-decoding algorithms, the target distribution is preserved through verification and sampling. The smaller model proposes candidates; the target model determines whether those candidates can become part of the generated sequence.
That distinguishes speculative decoding from simply replacing the target model with a smaller, faster model. The purpose is to change the execution path, not to accept lower-quality predictions as the final answer.
“Lossless” does not mean every implementation on every piece of hardware will produce byte-for-byte identical floating-point behavior. It refers to preserving the target model’s sampling distribution algorithmically, subject to numerical implementation details. It also says nothing about whether speculation will actually make a particular workload faster.
Speculation spends compute to save sequential time
This is the part that makes speculative decoding an infrastructure story rather than just a sampling trick.
During low-batch decoding, the target model can be limited by moving model data through the memory hierarchy repeatedly. The GPU may have arithmetic capability that the workload cannot fully exploit. Verification of several candidate positions gives the accelerator more parallel work per expensive target-model invocation.
That is the opportunity. Extra speculative computation can be attractive when compute is available but sequential target-model execution is the bottleneck.
The balance changes when the server becomes busy.
A GPU serving many concurrent sequences already has more real work available. Its spare compute headroom can shrink. Candidate tokens then compete with actual user tokens for the same physical accelerator, and rejected candidates become more expensive because the compute they consumed could have advanced other requests.
The same optimization can behave differently under load
Imagine one interactive request on a large GPU. Decode is relatively narrow, the accelerator has spare compute, and reducing expensive sequential target-model steps can directly improve responsiveness.
Now imagine hundreds of active sequences. The serving engine can already batch useful work from many requests. The GPU is no longer waiting for enough arithmetic to do.
At that point speculative candidates are not filling empty capacity in the same way. They are joining the competition for it.
This connects speculative decoding to the broader capacity problem described in How Many Users Can One GPU Actually Serve? An optimization that improves one user’s generation latency is not automatically an optimization for maximum multi-user throughput.
Compare speculative and ordinary decoding at the same model, hardware, prompts and concurrency. Track inter-token latency and end-to-end latency, but also aggregate throughput, GPU utilization and the proportion of speculative candidates that are accepted. A faster-looking single request can hide a worse server-level trade-off.
There is more than one way to speculate
The small-draft-model design is the easiest mechanism to visualize, but modern serving systems support several approaches.
vLLM currently documents model-based techniques including EAGLE-style speculators, multi-token prediction, draft models and other proposers, alongside simpler approaches such as n-gram and suffix speculation.
The physical objective remains similar: produce useful candidate future work more cheaply than repeatedly invoking the full target path one token at a time.
But the cost profile changes. A separate draft model needs its own execution and potentially additional memory. Other methods can reuse information differently. The correct choice therefore depends on model support, workload, memory headroom and the latency-versus-throughput objective.
| Observation | What it may mean | What to inspect |
|---|---|---|
| Latency improves at low concurrency | Spare compute is being traded effectively for fewer sequential steps | ITL, accepted candidates, GPU utilization |
| Gain disappears as load rises | Speculative work is competing with useful batched work | Concurrency, throughput, GPU saturation |
| Many candidates are rejected | The proposer is doing work that rarely advances output | Acceptance statistics and proposer configuration |
| VRAM pressure increases | The speculative method has added model or runtime state | Model residency, KV cache and runtime allocations |
A useful first-party experiment would compare ordinary and speculative decoding on one fixed GPU while gradually increasing concurrency. The interesting point is not the maximum speedup at one request, but where speculative work stops being cheap as the accelerator fills with real serving work.
The fastest token may be the target-model step you avoid
Speculative decoding does not make autoregressive generation stop being autoregressive. The final sequence still has to satisfy the target model.
What changes is how the hardware reaches that sequence. Instead of forcing the expensive model to discover every future token through a separate narrow step, the system spends cheaper work proposing a short path ahead and asks the target model to validate more of that path at once.
On an underfilled GPU, that trade can turn spare compute into lower latency. On a saturated server, the same speculative work can become another consumer of a resource that is no longer spare.
So the useful question is not whether speculative decoding is faster. It is whether your GPU has the right kind of unused capacity for speculation to buy back expensive sequential time.







