Inference
Why a Local LLM Is Slow: How to Find the Real Inference Bottleneck
1137
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work