CPU Offloading
A local LLM does not necessarily have to fit entirely inside your GPU’s VRAM. If the model is too large, runtimes such as llama.cpp can keep part
Local AI, Local LLMs, GPU hardware, inference, RAG and AI server infrastructure explained clearly.