Servers
Speculative Decoding Explained: How LLMs Generate More Than One Token per Expensive Step
073
A large language model can have a powerful GPU almost to itself and still produce a single conversation one token at a time. The accelerator finishes one
Inference
Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers
4152
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.