Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers

Inference

A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt. After that, it generates new tokens one at a time. These are different workloads, and the same computer can be fast at one while being much slower at the other.

This distinction explains a common local-AI experience: a short chat feels responsive, but pasting a long document creates a noticeable pause before anything appears. It also explains why two benchmark results for the same model and GPU can show dramatically different token rates without contradicting each other.

The useful measurements are prompt processing speed and token generation speed. Understanding what each represents makes it much easier to diagnose a slow local LLM and decide whether changing the model, context, runtime, or hardware is likely to help.

Local LLM speed has two main phases

For a typical text-generation request, inference can be simplified into two computational phases:

  1. Prompt processing — the model processes the input tokens that already exist.
  2. Token generation — the model predicts new output tokens sequentially.

In LLM literature and software, you may also see the first phase called prefill and the second called decode. llama.cpp uses the terms prompt processing and text generation in its benchmarking tools.

This is not merely a difference in terminology. The two phases use the model differently.

Phase What it does Typical user experience Useful metric
Prompt processing Processes existing input tokens Waiting before the answer begins Prompt tokens/s
Token generation Produces new tokens Speed at which the answer appears Generated tokens/s

There are other contributors to latency, including model loading, tokenization, sampling, request scheduling and application overhead. But for an already loaded model performing ordinary inference, prompt processing and generation are the two numbers that explain much of what the user perceives as model speed.

What happens during prompt processing?

Suppose you send a local model a prompt containing 4,000 tokens. Those tokens might include a system prompt, previous conversation turns, retrieved RAG passages and your latest question.

Before the model can produce a useful next token, it needs to run that input through the model and establish the internal state needed for generation. During this phase, the runtime can process multiple prompt tokens together rather than treating the entire prompt like sequential output generation.

That ability to process prompt tokens in batches is one reason prompt-processing throughput can be much higher than generation throughput on suitable hardware.

The official llama-bench documentation explicitly separates the workloads. Its -p option benchmarks prompt processing, while -n benchmarks text generation. It also exposes batch-size controls for prompt processing. The llama.cpp benchmark documentation describes these as separate test types.

FACT: prompt processing can use batching, and its throughput is sensitive to parameters such as prompt length, batch configuration, model, backend and hardware.

That means a prompt-processing result should not be interpreted as “this model generates text at this speed.” It measures how quickly the runtime can consume existing input under the tested configuration.

What happens during token generation?

Generation has a different dependency.

After the prompt has been processed, the model predicts a next token. That token becomes part of the context. The model then predicts another token based on the updated context, followed by another, and so on.

The important part is that these output tokens are sequential. Token 101 depends on the state produced after token 100, so a single response cannot simply generate hundreds of future tokens independently in one large batch.

This makes generation fundamentally different from processing a block of already-known prompt tokens.

If your model generates at a comfortable rate, text appears smoothly after the initial wait. If generation is slow, you may get the first token quickly but then watch the response arrive painfully slowly.

In llama.cpp timing output, these phases are reported separately as prompt eval time and eval time. The server code also tracks separate prompt-processing and generation token rates. llama.cpp’s server implementation exposes this distinction directly in its timing statistics.

Why prompt processing can be much faster than generation

It is tempting to assume that if the same model and weights are involved, both phases should have roughly the same throughput. They often do not.

Prompt processing has more opportunities to execute work across multiple known tokens together. This can make better use of highly parallel compute resources such as GPUs.

Generation of one sequence has a stronger sequential dependency: the next output token is not known until the preceding step has completed.

As a result, the performance characteristics of the hardware can matter differently in each phase.

A powerful GPU may achieve very high prompt-processing throughput because it can perform substantial parallel work. Generation can still run at a much lower tokens-per-second rate because the workload cannot exploit the same degree of token-level parallelism for one sequence.

This is why benchmark numbers need labels. “500 tokens/s” is not useful by itself if you do not know whether it refers to prompt processing, generation, or aggregate throughput across multiple concurrent requests.

Prompt speed affects time to first token

The difference becomes especially visible with long inputs.

Consider two requests using the same model:

  • a short question containing a few dozen input tokens;
  • a document-analysis request containing thousands of input tokens.

The generation speed after the answer starts may be similar in both cases. But the second request gives the model far more input to process first.

This contributes to time to first token: how long the user waits before output begins.

Prompt processing is not the only contributor to that latency. A cold request may also need to load the model, and a busy local server may queue the request. Tokenization and application overhead add additional time. Prompt caching can reduce the amount of input that needs to be recomputed in some runtimes.

Still, once those factors are controlled, increasing the amount of uncached input generally means more prompt-processing work.

This matters for local workflows that routinely send large inputs: RAG, document summarization, code repositories, long chat histories and agent systems with large tool outputs.

Generation speed affects how fast the answer feels

Once output begins, generation throughput becomes the more visible metric.

If you mostly use a local LLM for interactive chat, coding assistance or short questions, generation speed may dominate your perception of performance because you spend most of the interaction watching new tokens arrive.

If you use the model for classification where the output is only a few tokens, generation throughput may matter much less. Processing the input may account for a larger share of the request.

The workload therefore determines which performance number deserves more attention.

Workload Likely important phase Why
Interactive chat Both Users notice both initial latency and streaming speed
Long-document Q&A Prompt processing Large input must be processed before useful output begins
Long-form generation Token generation A large number of output tokens must be decoded sequentially
Classification Prompt processing Input may be much larger than the output
RAG Often both Retrieved passages enlarge the prompt, while the final answer still requires generation

How to read Ollama performance numbers

Ollama exposes enough timing information through its API to separate these phases instead of relying on how fast the interface feels.

According to the Ollama Generate API documentation, a completed response can include:

  • load_duration — time spent loading the model;
  • prompt_eval_count — number of input tokens;
  • prompt_eval_cached_count — prompt tokens read from cache;
  • prompt_eval_duration — time spent evaluating uncached prompt tokens;
  • eval_count — number of generated output tokens;
  • eval_duration — time spent generating them.

You can derive approximate throughput from these counters:

prompt processing tokens/s = uncached prompt tokens / prompt evaluation time

generation tokens/s = generated tokens / generation time

Ollama reports the durations in nanoseconds, so the duration must be converted to seconds before calculating tokens per second.

RECOMMENDATION: when troubleshooting Ollama performance, record both prompt and generation statistics. Do not collapse them into one average tokens-per-second figure.

How to benchmark both phases with llama.cpp

For llama.cpp, llama-bench is useful because it deliberately separates prompt processing from text generation.

A basic benchmark can vary the number of prompt tokens with -p and generated tokens with -n. The tool also supports a combined prompt-processing-plus-generation test.

FACT: the llama.cpp documentation notes that llama-bench measurements do not include tokenization and sampling time. That is useful for comparing core inference performance, but it also means the benchmark is not identical to complete end-to-end application latency.

For meaningful comparisons, keep the important variables fixed: model file, quantization, GPU offload, context configuration, batch settings, backend and hardware.

Changing several of them at once makes it difficult to identify why performance changed.

Why batch size changes prompt performance

Prompt-processing throughput can depend substantially on batching.

llama.cpp exposes both batch and micro-batch controls, and its benchmark documentation specifically demonstrates testing prompt processing with different batch sizes.

A larger batch can give the backend more work to execute together, potentially improving hardware utilization. But larger is not automatically better. Batch configuration interacts with memory consumption, backend implementation, model architecture and the hardware itself.

RECOMMENDATION: do not copy a batch setting simply because it produced a high benchmark result on someone else’s machine. Benchmark the settings relevant to your own model and workload while watching memory use.

This is particularly important on systems close to their VRAM or unified-memory limit. A configuration that improves prompt throughput is not useful if it causes memory pressure, forces an undesirable fallback or prevents the model from running at the context size you actually need.

Context length can affect more than memory

Long context is often discussed as a memory problem, but it also creates computation.

Every additional uncached token supplied to a request has to be processed before it can influence the answer. Longer active sequences also increase the state the model must work with during generation.

The exact performance impact depends on model architecture, runtime implementation, attention optimizations, caching and hardware. It should not be reduced to a universal formula.

Practically, however, an oversized context can hurt local inference in two ways: it consumes additional memory and can add latency.

This is one reason setting the largest context window supported by a model is not automatically the best local configuration. The useful context size is the one your workload actually needs.

Prompt caching changes the picture

Not every request necessarily recomputes every input token.

When a runtime can reuse a compatible prefix from a previous computation, cached prompt state can reduce repeated prompt-processing work. This is particularly useful when many requests share a large system prompt or stable conversation prefix.

Ollama’s current API makes this visible by reporting prompt_eval_cached_count separately from the uncached prompt evaluation statistics.

This also means two apparently identical “4,000-token prompts” may have different latency if one can reuse cached state and the other must process the full input again.

FACT: cache behavior is runtime- and configuration-dependent. Do not assume that a long prompt is automatically cached simply because it was used before.

Hardware bottlenecks can look different in the two phases

If a local model feels slow, asking only “Is my GPU fast enough?” is too vague.

You first need to identify which phase is slow.

If prompt processing is the problem, relevant variables can include compute capability, batch configuration, how much of the model is accelerated, memory capacity and bandwidth, backend efficiency, and the size of the input.

If generation is the problem, model size, quantization, memory bandwidth, CPU/GPU placement and the cost of repeatedly running the model for each new token become especially important.

CPU offloading complicates both phases further. If some model work occurs on the CPU while other work runs on the GPU, data movement and the relative performance of both devices can affect the result. The exact penalty depends on the runtime, model, offload configuration and machine.

There is therefore no trustworthy universal statement such as “this GPU runs this model at X tokens/s” without defining the test.

Do not confuse single-user generation with server throughput

There is another important source of misleading performance numbers: concurrency.

A local server handling several sequences can batch work across requests. That can increase aggregate throughput even though an individual user’s sequence is still generated autoregressively.

For example, a benchmark may report the combined number of tokens generated for several parallel sequences per second. That is a useful server-capacity metric, but it is not the same thing as the rate at which one chat window displays tokens.

The llama.cpp batched benchmark explicitly reports prompt-processing time, generation time and aggregate throughput for batches, illustrating why throughput and per-request responsiveness should be treated as separate measurements.

When comparing benchmark results, always ask:

  • Is this prompt processing or generation?
  • Is the result for one sequence or several?
  • What prompt and output lengths were tested?
  • Which quantization and model file were used?
  • How much of the model was GPU-offloaded?
  • What batch and context settings were used?

How to diagnose a local LLM that feels slow

A useful diagnosis does not start by changing random runtime parameters. Start by separating the delay into stages.

1. Check model loading

If only the first request is slow and later requests are responsive, model loading may account for much of the delay. Do not mistake load time for inference performance.

2. Check time before the first token

If the model is already loaded but long prompts create a large pause, inspect prompt-processing statistics. Compare a short prompt with a representative long prompt.

3. Check generation speed

Once the first token appears, watch how quickly subsequent tokens arrive. If this phase is slow even with a tiny prompt, the bottleneck is unlikely to be simply the amount of input text.

4. Compare like with like

Keep the model, quantization, runtime and offload configuration fixed while changing one variable at a time. Otherwise a “faster” result may simply be measuring a different workload.

5. Test the workload you actually run

A 512-token synthetic prompt is useful for repeatable benchmarking, but it may say little about a RAG pipeline that routinely submits 12,000 tokens. Likewise, a prompt-processing benchmark alone does not tell you how pleasant long-form generation will feel.

Which speed number should you optimize?

There is no single answer because the correct target follows from the workload.

For an interactive assistant, both time to first token and generation rate matter. For document classification, prompt throughput may be more important. For long-form writing, generation throughput can dominate total runtime. For RAG, unnecessarily large retrieved contexts can make prompt processing an avoidable source of latency.

RECOMMENDATION: treat local LLM performance as at least two measurements rather than one:

  • How quickly can the system process the input?
  • How quickly can it generate the output?

Then add model-loading time and concurrency metrics when they matter to your deployment.

This produces a much more useful performance picture than a single tokens-per-second number. More importantly, it tells you what to change. A long wait before the first token and slow text generation are different problems, even when they happen on the same model, on the same computer, in the same request.

Rate article
Add a comment