Q4_K_M
A 14 GB model file does not mean you need exactly 14 GB of memory to run it. And a model described as “8B” does not have one fixed memory requirement.
Local AI, Local LLMs, GPU hardware, inference, RAG and AI server infrastructure explained clearly.