Local LLMs
How LLM Quantization Works: Q4, Q5, Q8 and What You Should Actually Use
0103
A 14 GB model file does not mean you need exactly 14 GB of memory to run it. And a model described as “8B” does not have one fixed memory requirement.