Inference
Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers
4152
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.
Inference
Why a Local LLM Is Slow: How to Find the Real Inference Bottleneck
1136
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work
Hardware
Context Windows and VRAM: Why Longer Conversations Need More Memory
2117
A local LLM can fit comfortably into your GPU at an 8,000-token context and run out of memory when you increase that context several times over.