Inference
Your LLM Benchmark May Be Measuring the Wrong Thing
1116
A local LLM produces 100 tokens per second. On paper, that sounds fast. Then a real user sends a long prompt and waits several seconds before anything