LM Studio GGUF VRAM & Layer Calculator
Calculate required VRAM or Apple Unified Memory for running open-weights LLMs locally with GGUF quantizations in LM Studio.
What these numbers mean
The footprint figure is the model's weights plus the KV cache plus a small allowance for application overhead. Weight size scales with parameter count and bits per weight, so a 14B model at Q4_K_M is roughly half the size of the same model at Q8_0. The KV cache is the conversation's working memory: it grows with context length, and it is the term most people forget when a model that "just fits" turns out to run badly.
Treat the result as a sizing estimate, not a guarantee. Exact KV cache size depends on a model's layer count and attention layout, and quantized KV cache or Flash Attention will move the figure. The rule that matters is the one the estimate is testing: if the total exceeds your fast memory, layers spill to the CPU and generation speed collapses toward the slowest tier rather than degrading gently.
- Choosing a GGUF quantization level — what Q4_K_M actually costs you against Q8_0.
- LM Studio system requirements — the hard platform floors and how to size a machine.
- GGUF vs safetensors — why a model may not load at all, regardless of memory.
- LM Studio on Unraid — running the model on a server instead of your desktop.