Determining VRAM Requirements for Local LLM Execution

The primary consideration when deploying a local LLM is ensuring it fits within your GPU's memory. The outcome hinges on the model's size, the level of quantization, and the context length. This guide offers a practical framework for selecting the appropriate VRAM capacity.

The impact of quantization on VRAM

Quantization reduces the precision used to store model weights. Lower-bit quantization decreases the model's footprint and VRAM consumption, albeit with a slight trade-off in quality.

Quant Bits per weight Typical application
Q8_0 8 Premium quality
Q6_K ~6.6 High quality
Q5_K_M ~5.5 Optimal balance of quality and size
Q4_K_M ~4.5 Efficient balance of size and performance
Q3_K_M ~3.5 Minimal VRAM, higher quality loss

Q4_K_M is frequently selected when VRAM is constrained. If you have surplus VRAM, opting for Q5 or Q6 allows you to run the same model with reduced quantization.

Estimated VRAM based on model size

The following figures are approximate estimates for model weights only. The total VRAM requirement will be higher, as the runtime, KV cache, and context also consume memory.

Model size Q8_0 Q6_K Q4_K_M Q3_K_M
4B ~5 GB ~4 GB ~3 GB ~2.5 GB
8B ~9 GB ~7 GB ~5.5 GB ~4.5 GB
12B ~13 GB ~10 GB ~8 GB ~6.5 GB
14B ~16 GB ~12 GB ~9 GB ~7.5 GB
27B ~30 GB ~22 GB ~17 GB ~13 GB
32B ~36 GB ~27 GB ~20 GB ~16 GB
70B ~80 GB ~60 GB ~42 GB ~34 GB

These figures are estimates rather than strict limits. Variations in model architecture and quantization formats can alter the actual size.

Capabilities based on VRAM capacity

VRAM Practical range Current examples
8 GB Compact models ranging from 4B to 9B Gemma 4 E4B, Qwen3.5 9B
12 GB Compact to mid-sized models ranging from 9B to 14B Gemma 4 12B, Qwen3.5 9B
16 GB 12B to 27B with lower quantization Gemma 4 26B-A4B, Qwen3.6 27B at Q4
24 GB 27B to 35B at Q4 to Q6 Qwen3.8 27B, Gemma 4 31B
32 GB 27B to 35B at higher quantization Qwen3.8 27B, Gemma 4 31B
48 GB Large dense models at lower quantization 70B-class models at Q3 to Q4
80 GB Large dense models at higher quantization 70B-class models at Q4 to Q6

These ranges apply to models whose weights can reside entirely on the GPU. Large MoE models operate differently: while only a subset of parameters is active per token, the model must still store its entire weight set. Consequently, a model with 100B or more total parameters will not fit in a 100B-sized VRAM budget simply due to having fewer active parameters.

MoE models

Mixture-of-Experts models comprise multiple parameter groups known as experts. Since only specific experts are utilized for each token, inference can be more efficient compared to a dense model with the same total parameter count.

Nevertheless, inactive experts remain part of the model. Thus, large MoE models may demand significantly more memory than their active parameter count implies. Very large models might necessitate multiple GPUs or offloading to system RAM.

Context length consumes additional VRAM

Model weights represent only a portion of the memory requirement. The KV cache expands as the context lengthens, meaning running the same model with a 64K context can demand substantially more VRAM than a 4K context.

  • Extended context requires additional VRAM.
  • KV-cache precision impacts memory consumption.
  • Batch size and concurrent users also elevate memory usage.
  • Reserve VRAM for the runtime instead of filling the GPU exclusively with model weights.

Practical recommendations

  • Verify the actual size of the quantized model you intend to run.
  • Avoid using the model file size as the precise VRAM requirement. Allow space for the KV cache and runtime.
  • If a model does not fully fit in VRAM, some parts can be offloaded to system RAM, though inference speed typically decreases.
  • For long-context or agentic workloads, allocate more VRAM than what is required for the model weights alone.
  • Multiple GPUs can distribute a model if a single GPU lacks sufficient VRAM.

Execute on DaDesktop

There is no need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop with the necessary VRAM, allowing you to run the model directly without owning the hardware.

Select the VRAM tier that suits your model, load it, and begin using it. Enjoy a seamless experience with no setup, no hardware purchase, and no driver issues. See available GPUs for your options.