Servers
Why Memory Bandwidth Matters More Than CPU Core Count for Local LLMs.
016
A 16-core CPU can load a quantized local LLM, show activity across many cores, and still generate text much more slowly than expected. Adding threads may help at first.
Inference
Why a Local LLM Is Slow: How to Find the Real Inference Bottleneck
1136
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work