Servers
Why Memory Bandwidth Matters More Than CPU Core Count for Local LLMs.
016
A 16-core CPU can load a quantized local LLM, show activity across many cores, and still generate text much more slowly than expected. Adding threads may help at first.
Local LLMs
How LLM Quantization Works: Q4, Q5, Q8 and What You Should Actually Use
0102
A 14 GB model file does not mean you need exactly 14 GB of memory to run it. And a model described as “8B” does not have one fixed memory requirement.
Hardware
Context Windows and VRAM: Why Longer Conversations Need More Memory
2117
A local LLM can fit comfortably into your GPU at an 8,000-token context and run out of memory when you increase that context several times over.