Servers
Continuous Batching Explained: How LLM Servers Keep a GPU Busy
168
Four people are talking to the same local LLM. One asks a short question and gets an answer in seconds. Another requests a long explanation.
Inference
Your LLM Benchmark May Be Measuring the Wrong Thing
1116
A local LLM produces 100 tokens per second. On paper, that sounds fast. Then a real user sends a long prompt and waits several seconds before anything
Inference
Prompt Processing vs Token Generation: Why Local LLM Speed Has Two Numbers
4152
A local LLM does not have one meaningful “tokens per second” number. Before it can write the first token of an answer, it has to process the tokens already in the prompt.
Hardware
CPU Offloading Explained: How to Run a Local LLM That Doesn’t Fit in VRAM
0146
A local LLM does not necessarily have to fit entirely inside your GPU’s VRAM. If the model is too large, runtimes such as llama.cpp can keep part
Local LLMs
How Much Context Should You Give a Local LLM?
0148
A local model may advertise a 32K, 64K, 128K, or even larger context window. That does not mean you should configure your local runtime to use the maximum.
Inference
Why a Local LLM Is Slow: How to Find the Real Inference Bottleneck
1136
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work
Hardware
How Much VRAM Do You Actually Need to Run an LLM Locally?
182
A 14 GB model file does not mean you need exactly 14 GB of VRAM to run it. The weights are only one part of the memory used during inference.