Tools
Interactive tools built and maintained by LM Studio Guide. Free, no signup, and everything runs in your browser, so nothing about your hardware or your models is sent anywhere.
Estimate the memory a model needs before you download 40 GB of weights. Choose a parameter count, a GGUF quantization level, a context window, and a hardware target, and it returns the total memory footprint, the KV cache portion, and a suitable GPU or unified-memory tier.
Open the sizer
What it is for
The single constraint that decides whether a local model is pleasant or unusable is whether it fits in fast memory. A model that overflows VRAM does not slow down a little; the layers that spill to the CPU become the bottleneck for every token, and throughput falls toward the speed of the slowest tier. The sizer exists to answer that question in advance rather than after a long download.
Two variables move the answer most. Quantization level sets the size of the weights, roughly parameter count multiplied by bits per weight. Context length sets the size of the KV cache, which is the term people routinely forget when a model that appeared to fit turns out to leave no room for an actual conversation.
Guides that go with it
- GGUF quantization levels: Q4_K_M vs Q8_0 — which quantization to pick and what it costs.
- LM Studio system requirements — the hard platform floors, and sizing a machine to a model.
- GGUF vs safetensors — which formats LM Studio can open, and what to do when a model will not load.
- LM Studio on Unraid — headless and containerised options for serving models from a home server.