GPU offloading
A local LLM being slow does not automatically mean that your GPU is too weak. The model may be spilling into system RAM, the CPU may be doing more work
Local AI, Local LLMs, GPU hardware, inference, RAG and AI server infrastructure explained clearly.