Servers
Continuous Batching Explained: How LLM Servers Keep a GPU Busy
167
Four people are talking to the same local LLM. One asks a short question and gets an answer in seconds. Another requests a long explanation.
Inference
How Many Users Can One GPU Actually Serve?
198
A local AI server can feel almost instantaneous when one person is using it. Add a second conversation and it may still feel fine. Add several people with
Inference
Your LLM Benchmark May Be Measuring the Wrong Thing
1115
A local LLM produces 100 tokens per second. On paper, that sounds fast. Then a real user sends a long prompt and waits several seconds before anything