LLM System Requirements Calculator
Estimate the VRAM, system RAM, and disk space you need to run any open-source LLM locally. Covers Llama, Mistral, Qwen, DeepSeek, Gemma, Phi and 30+ models across FP16 through Q2_K quantization.
- Free, no sign-up
- REST + MCP
- Updated
- Reviewed by Olgun Ozoktas
Your System
Configure your model
Apache 2.0 · released 2024-05 · native context up to 32K tokens · HuggingFace
Standard community default — best size/quality tradeoff
Typical: 4K–8K chat, 32K–128K long-doc work. Longer context = larger KV cache.
Will it run on my hardware?
| Hardware | Memory | Fit | Headroom | Note |
|---|---|---|---|---|
| NVIDIA RTX 3060 12GB | 12 GB | Fits | +2.58 GB | Budget consumer GPU |
| NVIDIA RTX 4060 Ti 16GB | 16 GB | Fits | +6.58 GB | Mid-range with extra VRAM |
| NVIDIA RTX 3090 24GB | 24 GB | Fits | +14.6 GB | Enthusiast, great for local LLMs |
| NVIDIA RTX 4090 24GB | 24 GB | Fits | +14.6 GB | Top consumer GPU |
| NVIDIA RTX 5090 32GB | 32 GB | Fits | +22.6 GB | Current-gen flagship |
| NVIDIA A100 40GB | 40 GB | Fits | +30.6 GB | Datacenter-class |
| NVIDIA A100 80GB | 80 GB | Fits | +70.6 GB | Datacenter-class, large |
| NVIDIA H100 80GB | 80 GB | Fits | +70.6 GB | Current datacenter flagship |
| Mac M2 16GB | 12 GB(unified) | Fits | +2.58 GB | Unified memory; ~12GB usable after OS |
| Mac M3 Max 36GB | 30 GB(unified) | Fits | +20.6 GB | Unified memory; ~30GB usable |
| Mac M3 Max 64GB | 56 GB(unified) | Fits | +46.6 GB | Unified memory; ~56GB usable |
| Mac M3 Ultra 128GB | 115 GB(unified) | Fits | +105.6 GB | Unified memory; ~115GB usable |
| Mac M4 Max 128GB | 115 GB(unified) | Fits | +105.6 GB | Unified memory; ~115GB usable |
How the math works:
- Weights: total params × bytes/param. MoE routing reduces per-token compute, but normal resident memory still includes all expert weights.
- KV cache: a conservative FP16 estimate that scales with context and batch size. Cache precision and attention layout can reduce it.
- Overhead: activations, framework workspace, CUDA kernels — typically 1.5–3 GB or ~10% of weights, whichever is larger.
- Disk: the full quantized checkpoint. MoE models store every expert on disk even if only a few are active per token.
Sources: Architecture numbers (num_hidden_layers, hidden_size, max_position_embeddings) come directly from each model's published HuggingFace config.json — click the HuggingFace link next to the selected model to see the exact config. Quantization byte/param ratios come from the llama.cpp GGUF k-quants spec. Treat the result as a planning estimate because cache precision, offload, and framework buffers differ by runtime.
Quantization availability: Mathematically any model can be quantized to any level — but community GGUF releases don't always ship every variant. Q4_K_M, Q5_K_M, Q8_0 and FP16 are near-universal; Q2_K and Q3_K_M are often skipped on smaller models where quality loss is noticeable. To find a specific quant for a specific model, search bartowski, TheBloke, or the official repo on HuggingFace.
Context length: Each model has a native max context defined in its config. Going beyond it (via RoPE/YaRN scaling) works mechanically but degrades quality — the warning banner above flags this whenever the current selection exceeds the model's native limit.
How to estimate LLM requirements
-
Pick the model
Select the open-source model you plan to run — Llama 3.1, Mistral, Qwen 2.5, DeepSeek R1, Gemma 2, Phi, or Code Llama. The dropdown is grouped by family and shows license and release date. -
Choose a quantization level
Q4_K_M is the community default: about 0.58 bytes per parameter and nearly indistinguishable from FP16 on most tasks. FP16 doubles the size but matches the original training precision. Q2_K is the smallest, with visible quality loss. -
Set the context length
Context length drives the KV cache, which scales linearly with tokens. A 128K-context session on Llama 3.1 70B adds over 10 GB to VRAM on top of the weights. Use the preset buttons for common sizes. -
Read the verdict
The three cards show estimated total VRAM, a system RAM planning value, and estimated download size. The hardware table compares each target with the same conservative estimate.
Who this is for
Choosing a GPU for local inference
Sizing a cloud instance
Picking a quantization for your hardware
Budgeting disk space
About this tool
This calculator estimates the VRAM, system RAM, and disk space required to run a selected open-weight language model on local hardware. It covers models across the Llama, Mistral, Qwen, DeepSeek, Gemma, Phi, and Code Llama families, with quantization choices from FP32 through Q2_K.
The estimate separates quantized model weights, a conservative KV-cache value, and runtime overhead. Model weight memory uses total parameter count and the selected bytes-per-parameter value. KV cache grows with context length and batch size. The final runtime can use more or less memory because attention layout, cache precision, GPU offload, framework buffers, and concurrent requests differ.
For MoE models such as Mixtral and DeepSeek V3, active parameters describe per-token compute. Normal resident memory still needs all expert weights unless the runtime explicitly offloads experts to system memory or storage. The calculator therefore uses total parameters for its default memory estimate and keeps active parameters as model information.
Pair this calculator with the AI Model Picker to choose a model, or use the linked model card to confirm architecture details before downloading weights.
How it compares
Memory estimator blog posts and Hugging Face Space demos exist, but most are model-specific, skip the KV cache entirely, or ignore runtime overhead — which is exactly the memory that makes a model OOM instead of OOM-near-miss. This calculator is model-agnostic across 30+ curated checkpoints, always includes KV cache and overhead, and lets you see the headroom on 13 specific hardware targets side by side.
The alternative — trying to load the model and watching nvidia-smi or macOS Activity Monitor — works but wastes a 40 GB download when the answer is no. This tool gives you the answer before you download anything, and runs entirely in your browser: the model list and math are static data, so your hardware choices and model preferences are never transmitted anywhere.
Tips for running local LLMs
- Q4_K_M is the community default for a reason — it produces the best quality per gigabyte on almost every model and saves roughly 75% vs FP16.
- On unified-memory Macs (M1/M2/M3/M4), the whole RAM pool doubles as VRAM, minus about 4 GB for macOS. A 36 GB M3 Max gives you ~30 GB usable for model weights.
- KV cache memory grows with context length and batch size. Some runtimes support lower-precision caches, so check the settings for your runtime.
- Mixture-of-Experts models route each token through only some experts, but a normal resident-memory estimate still includes all expert weights. Expert offloading is a separate runtime choice.
- For a 7B–8B model at 8K context, Q4_K_M fits comfortably in 8 GB of VRAM and runs well on a mid-range laptop GPU or a base-model M2 Mac.