Inference
LLM Prefix Caching Explained: Why Repeated Prompts Can Start Faster
075
Two requests can contain almost the same 20,000-token prompt and still reach the first generated token at very different times. On the first request, the
Inference
How Long Prompts Disrupt Shared LLM Inference
080
Three people are using the same local AI server. Two answers are already streaming at a comfortable pace. Then a third user pastes a long document and presses Send.
Inference
When Does a Second GPU Actually Help LLM Inference?
092
There is an appealingly simple upgrade path for a busy local AI server: install another GPU. Twice the accelerators should mean something close to twice
Inference
How Many Users Can One GPU Actually Serve?
196
A local AI server can feel almost instantaneous when one person is using it. Add a second conversation and it may still feel fine. Add several people with
Inference
Your LLM Benchmark May Be Measuring the Wrong Thing
1113
A local LLM produces 100 tokens per second. On paper, that sounds fast. Then a real user sends a long prompt and waits several seconds before anything
Inference
Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers
4149
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.
Inference
Why a Local LLM Is Slow: How to Find the Real Inference Bottleneck
1136
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work