Inference
Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers
4152
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.
Local LLMs
How Much Context Should You Give a Local LLM?
0148
A local model may advertise a 32K, 64K, 128K, or even larger context window. That does not mean you should configure your local runtime to use the maximum.