---
title: "LLM System Requirements Calculator"
description: "Estimate the VRAM, system RAM, and disk space you need to run any open-source LLM locally. Covers Llama, Mistral, Qwen, DeepSeek, Gemma, Phi and 30+ models across FP16 through Q2_K quantization."
url: https://findutils.com/developers/llm-requirements-calculator/
category: developers
---

# LLM System Requirements Calculator

Estimate the VRAM, system RAM, and disk space you need to run any open-source LLM locally. Covers Llama, Mistral, Qwen, DeepSeek, Gemma, Phi and 30+ models across FP16 through Q2_K quantization.

**Use this tool:** [LLM System Requirements Calculator](https://findutils.com/developers/llm-requirements-calculator/)

## Programmatic access

- REST id `llm-requirements-calculator`: POST https://api.findutils.com/api/tools/llm-requirements-calculator/execute (reference: https://findutils.com/api/llm-requirements-calculator/)
- MCP tool `llm_requirements_calculator` on https://mcp.findutils.com (reference: https://findutils.com/mcp/llm-requirements-calculator/)

## Tips for running local LLMs

- Q4_K_M is the community default for a reason — it produces the best quality per gigabyte on almost every model and saves roughly 75% vs FP16.
- On unified-memory Macs (M1/M2/M3/M4), the whole RAM pool doubles as VRAM, minus about 4 GB for macOS. A 36 GB M3 Max gives you ~30 GB usable for model weights.
- KV cache memory grows with context length and batch size. Some runtimes support lower-precision caches, so check the settings for your runtime.
- Mixture-of-Experts models route each token through only some experts, but a normal resident-memory estimate still includes all expert weights. Expert offloading is a separate runtime choice.
- For a 7B–8B model at 8K context, Q4_K_M fits comfortably in 8 GB of VRAM and runs well on a mid-range laptop GPU or a base-model M2 Mac.

## Frequently Asked Questions

### How much VRAM do I need to run Llama 3.1 70B?

At Q4_K_M with an 8K context, Llama 3.1 70B needs about 42 GB of VRAM (around 40 GB of weights plus 1 GB of KV cache and 3 GB of overhead). That fits a 2×3090 or 2×4090 setup, an A100 40GB with tight headroom, or an M3 Max with 64+ GB of unified memory. At FP16 it doubles to about 150 GB, requiring an H100 80GB pair or A100 80GB multi-GPU.

### What does Q4_K_M mean, and why is it the default?

Q4_K_M is a 4-bit group-wise quantization format from llama.cpp. It stores most weights in 4-bit blocks with per-block scale factors, reaching about 0.58 bytes per parameter. The 'K' means k-quants (improved block layout) and the 'M' means medium mix — important weights stay in higher precision. It's the default because quality loss is typically under 1% on reasoning benchmarks while size drops 75% vs FP16.

### Can I run DeepSeek V3 or DeepSeek R1 locally?

The full 671B MoE models are extremely large. Only part of the network is active for each token, but all expert weights must still reside somewhere. A consumer system normally needs a smaller distill model unless its runtime uses extensive CPU or storage offload.

### Why does my VRAM estimate jump when I increase context length?

The KV cache is 2 × num_layers × hidden_size × context_length × 2 bytes, and it sits in VRAM alongside the weights. On a 70B model with 80 layers and 8192 hidden size, each 1K of context adds ~2.5 GB. Going from 8K to 128K adds over 30 GB. Shortening context is usually the fastest way to reclaim VRAM without changing models.

### Is this calculator accurate for vLLM, TensorRT, and MLX?

It is a planning estimate, not a runtime guarantee. Cache precision, grouped-query attention, GPU offload, framework buffers, and concurrent sequences can change the final memory use. Confirm the result with your runtime documentation.

### Does the calculator handle Apple Silicon / M-series Macs?

Yes. Apple Silicon uses unified memory, so the model shares one pool with the operating system and other applications. The hardware table keeps operating-system headroom in its planning values.

### What's the difference between VRAM and RAM in the results?

VRAM is dedicated GPU memory. System RAM serves CPU inference and GPU offload. Unified-memory systems share one physical pool, but the operating system and applications also use it.

### Are these download sizes accurate for GGUF files?

The disk value estimates quantized tensor size from parameter count. Real GGUF files also contain metadata and can use mixed tensor types, so check the published file size before downloading.

## Related Tools

- [AI Model Picker](https://findutils.com/developers/ai-model-picker/)
- [AI Agent Starter Guide](https://findutils.com/ai-agent-starter-guide/)
- [Claude Code Usage Analyzer](https://findutils.com/developers/claude-code-usage-analyzer/)
- [dnpm Configurator](https://findutils.com/developers/dnpm-configurator/)
- [Cloudflare Cost Calculator](https://findutils.com/developers/cloudflare-cost-calculator/)
- [Chmod Calculator](https://findutils.com/developers/chmod-calculator/)
- [JWT Decoder](https://findutils.com/developers/jwt-decoder/)
- [cURL to Code](https://findutils.com/developers/curl-to-code/)
