Determining VRAM Requirements for Local LLM Execution
The primary consideration when deploying a local LLM is ensuring it fits within your GPU's memory. The outcome hinges on the model's size, the level of quantization, and the context length. This guide offers a practical framework for selecting the appropriate VRAM capacity.
The impact of quantization on VRAM
Quantization reduces the precision used to store model weights. Lower-bit quantization decreases the model's footprint and VRAM consumption, albeit with a slight trade-off in quality.
| Quant | Bits per weight | Typical application |
|---|---|---|
| Q8_0 | 8 | Premium quality |
| Q6_K | ~6.6 | High quality |
| Q5_K_M | ~5.5 | Optimal balance of quality and size |
| Q4_K_M | ~4.5 | Efficient balance of size and performance |
| Q3_K_M | ~3.5 | Minimal VRAM, higher quality loss |
Q4_K_M is frequently selected when VRAM is constrained. If you have surplus VRAM, opting for Q5 or Q6 allows you to run the same model with reduced quantization.
Estimated VRAM based on model size
The following figures are approximate estimates for model weights only. The total VRAM requirement will be higher, as the runtime, KV cache, and context also consume memory.
| Model size | Q8_0 | Q6_K | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 4B | ~5 GB | ~4 GB | ~3 GB | ~2.5 GB |
| 8B | ~9 GB | ~7 GB | ~5.5 GB | ~4.5 GB |
| 12B | ~13 GB | ~10 GB | ~8 GB | ~6.5 GB |
| 14B | ~16 GB | ~12 GB | ~9 GB | ~7.5 GB |
| 27B | ~30 GB | ~22 GB | ~17 GB | ~13 GB |
| 32B | ~36 GB | ~27 GB | ~20 GB | ~16 GB |
| 70B | ~80 GB | ~60 GB | ~42 GB | ~34 GB |
These figures are estimates rather than strict limits. Variations in model architecture and quantization formats can alter the actual size.
Capabilities based on VRAM capacity
| VRAM | Practical range | Current examples |
|---|---|---|
| 8 GB | Compact models ranging from 4B to 9B | Gemma 4 E4B, Qwen3.5 9B |
| 12 GB | Compact to mid-sized models ranging from 9B to 14B | Gemma 4 12B, Qwen3.5 9B |
| 16 GB | 12B to 27B with lower quantization | Gemma 4 26B-A4B, Qwen3.6 27B at Q4 |
| 24 GB | 27B to 35B at Q4 to Q6 | Qwen3.8 27B, Gemma 4 31B |
| 32 GB | 27B to 35B at higher quantization | Qwen3.8 27B, Gemma 4 31B |
| 48 GB | Large dense models at lower quantization | 70B-class models at Q3 to Q4 |
| 80 GB | Large dense models at higher quantization | 70B-class models at Q4 to Q6 |
These ranges apply to models whose weights can reside entirely on the GPU. Large MoE models operate differently: while only a subset of parameters is active per token, the model must still store its entire weight set. Consequently, a model with 100B or more total parameters will not fit in a 100B-sized VRAM budget simply due to having fewer active parameters.
MoE models
Mixture-of-Experts models comprise multiple parameter groups known as experts. Since only specific experts are utilized for each token, inference can be more efficient compared to a dense model with the same total parameter count.
Nevertheless, inactive experts remain part of the model. Thus, large MoE models may demand significantly more memory than their active parameter count implies. Very large models might necessitate multiple GPUs or offloading to system RAM.
Context length consumes additional VRAM
Model weights represent only a portion of the memory requirement. The KV cache expands as the context lengthens, meaning running the same model with a 64K context can demand substantially more VRAM than a 4K context.
- Extended context requires additional VRAM.
- KV-cache precision impacts memory consumption.
- Batch size and concurrent users also elevate memory usage.
- Reserve VRAM for the runtime instead of filling the GPU exclusively with model weights.
Practical recommendations
- Verify the actual size of the quantized model you intend to run.
- Avoid using the model file size as the precise VRAM requirement. Allow space for the KV cache and runtime.
- If a model does not fully fit in VRAM, some parts can be offloaded to system RAM, though inference speed typically decreases.
- For long-context or agentic workloads, allocate more VRAM than what is required for the model weights alone.
- Multiple GPUs can distribute a model if a single GPU lacks sufficient VRAM.
Execute on DaDesktop
There is no need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop with the necessary VRAM, allowing you to run the model directly without owning the hardware.
Select the VRAM tier that suits your model, load it, and begin using it. Enjoy a seamless experience with no setup, no hardware purchase, and no driver issues. See available GPUs for your options.