Skip to main content
Updated 2026-09-16 15 model-and-tier rows 5 model families Prices tracked live

Which GPU for Which LLM — VRAM Requirements by Model (2026)

A 32B model at Q4_K_M needs about 20 GB resident, which is why 24 GB is the floor for that class; 12 GB tops out at 14B and 70B wants 48 GB. The table below gives that mapping for Qwen3, Llama 4, DeepSeek V3.2, Gemma 4, Phi-4, with the VRAM floor, the throughput each tier is published at, and a live tracked price for a card that clears it. Per-card medians behind those figures come from the published runs aggregated on SpecPicks' benchmark pages.

As an Amazon Associate, SpecPicks earns from qualifying purchases. See our review methodology.

How to read the VRAM floor

Every floor in this guide comes from the same arithmetic rather than from a vendor chart: Q4_K_M weights run roughly 0.55 GB per billion parameters, and the runtime plus a usable context window wants about 2 GB on top. A 14B model is therefore ~9.7 GB resident, a 32B ~19.6 GB, and a 70B ~40 GB before context. Round up to the nearest card capacity and that is the floor.

The figure that stops being predictable is what happens below the floor. Once weights spill past VRAM the number is set by PCIe and system-memory bandwidth, not by the GPU, and the drop is steep rather than gradual — which is why a tier is quoted as a floor and not as a suggestion. Per-card medians over published runs, with the run counts and sources behind them, live on the benchmark pages; the GPU picker ranks cards that clear a floor by price.

Qwen3 — VRAM floors and GPU picks

Variants 1.7B · 4B · 7B · 14B · 32B · 72B · quantisation Q4_K_M · context 128K · full Qwen3 build guide →

TierVRAM floorPublished throughputCard to buy at this tier
Qwen3 🧪 Minimum (7B/14B at Q4) 12 GB32 GB DDR5 40–80 tok/s on 14B Q4
ZOTAC Gaming GeForce RTX 4060 Ti 16GB AMP DLSS 3 16GB… Tracked at $550 on 2026-09-16 — price may vary. View Current Price →
Qwen3 🏆 Recommended (32B at Q4) 24 GB64 GB DDR5 ~25 tok/s on 32B Q4 (RTX 4090) · ~38 tok/s (RTX 5090 with 32B at higher quant)
msi Gaming GeForce RTX 4090 24GB GDRR6X 384-Bit HDMI/DP… Tracked at $2,950 on 2026-09-16 — price may vary. View Current Price →
Qwen3 ⚡ Luxury (72B at Q4) 48 GB128 GB DDR5 ~14 tok/s on 72B Q4 (dual RTX 3090 NVLink)
No tracked listing clears 48 GB in the catalogue today. Rank every card that does →

Perfect for chatbot + coding use cases on smaller variants. 16GB lets you run 14B at Q4 with comfortable context.

Llama 4 — VRAM floors and GPU picks

Variants 8B · 17B · 70B · 405B · quantisation Q4_K_M · context 256K · full Llama 4 build guide →

TierVRAM floorPublished throughputCard to buy at this tier
Llama 4 🧪 Minimum (8B at Q4) 8 GB32 GB DDR5 ~95 tok/s on 8B Q4 (RTX 4070)
ZOTAC Gaming GeForce RTX 4060 Ti 16GB AMP DLSS 3 16GB… Tracked at $550 on 2026-09-16 — price may vary. View Current Price →
Llama 4 🏆 Recommended (17B / 70B at Q4) 24 GB64 GB DDR5 ~45 tok/s on 17B (RTX 4090) · ~14 tok/s on 70B Q4 with offload (RTX 5090)
msi Gaming GeForce RTX 4090 24GB GDRR6X 384-Bit HDMI/DP… Tracked at $2,950 on 2026-09-16 — price may vary. View Current Price →
Llama 4 ⚡ Luxury (70B native, 405B with offload) 48 GB256 GB DDR5 ~22 tok/s on 70B Q4 (dual RTX 3090) · ~3 tok/s on 405B Q4 (Mac M3 Ultra 192 GB)
No tracked listing clears 48 GB in the catalogue today. Rank every card that does →

8B Q4 fits in 8 GB but you want 12-16 GB for context room. 8B is a good lightweight assistant; coding and complex reasoning are limited.

DeepSeek V3.2 — VRAM floors and GPU picks

Variants 16B (active 2.4B) · 236B MoE (active 21B) · quantisation Q4_K_M · context 128K · full DeepSeek V3.2 build guide →

TierVRAM floorPublished throughputCard to buy at this tier
DeepSeek V3.2 🧪 Minimum (16B dense at Q4) 16 GB32 GB DDR5 ~60 tok/s on 16B Q4 (RTX 4070 Ti Super 16 GB)
ZOTAC Gaming GeForce RTX 4060 Ti 16GB AMP DLSS 3 16GB… Tracked at $550 on 2026-09-16 — price may vary. View Current Price →
DeepSeek V3.2 🏆 Recommended (16B at Q4 with full 128K context) 24 GB64 GB DDR5 ~75 tok/s on 16B Q4 (RTX 4090) · ~95 tok/s (RTX 5090)
msi Gaming GeForce RTX 4090 24GB GDRR6X 384-Bit HDMI/DP… Tracked at $2,950 on 2026-09-16 — price may vary. View Current Price →
DeepSeek V3.2 ⚡ Luxury (236B MoE) 96 GB256 GB DDR5 ~12 tok/s on 236B Q4 (Mac M3 Ultra) · ~18 tok/s (4× RTX 4090)
No tracked listing clears 96 GB in the catalogue today. Rank every card that does →

16 GB cards run the 16B-dense variant at Q4 cleanly. Avoid 8 GB cards — context window collapses.

Gemma 4 — VRAM floors and GPU picks

Variants 2B · 7B · 9B · 27B · quantisation Q4_K_M · context 128K · full Gemma 4 build guide →

TierVRAM floorPublished throughputCard to buy at this tier
Gemma 4 🧪 Minimum (2B / 7B at Q4) 8 GB16 GB DDR5 ~140 tok/s on 2B Q4 · ~75 tok/s on 7B Q4 (RTX 4060)
msi Gaming GeForce RTX 4060 Ti 8GB GDRR6 Extreme Clock… Tracked at $488 on 2026-08-24 — price may vary. View Current Price →
Gemma 4 🏆 Recommended (9B / 27B at Q4) 16 GB32 GB DDR5 ~85 tok/s on 9B Q4 (RTX 4070 Ti Super 16 GB) · ~35 tok/s on 27B Q4 (RTX 4090)
Gemma 4 ⚡ Luxury (27B at Q8 with 128K context) 24 GB64 GB DDR5 ~28 tok/s on 27B Q8 (RTX 4090) · ~42 tok/s (RTX 5090)
msi Gaming GeForce RTX 4090 24GB GDRR6X 384-Bit HDMI/DP… Tracked at $2,950 on 2026-09-16 — price may vary. View Current Price →

Gemma 4 2B at Q4 fits even on integrated graphics. 7B Q4 needs 8 GB; 12 GB lets you keep full 128K context active.

Phi-4 — VRAM floors and GPU picks

Variants mini (4B) · 14B · quantisation Q4_K_M · context 16K · full Phi-4 build guide →

TierVRAM floorPublished throughputCard to buy at this tier
Phi-4 🧪 Minimum (mini at Q4) 6 GB16 GB DDR5 ~180 tok/s on Phi-4 mini Q4 (RTX 4060) · ~90 tok/s on RTX 3060 12 GB
msi Gaming GeForce RTX 4060 Ti 8GB GDRR6 Extreme Clock… Tracked at $488 on 2026-08-24 — price may vary. View Current Price →
Phi-4 🏆 Recommended (14B at Q4 or Q8) 16 GB32 GB DDR5 ~95 tok/s on 14B Q4 (RTX 4070 Ti Super 16 GB) · ~62 tok/s on 14B Q8 (RTX 4090)
Phi-4 ⚡ Luxury (14B FP16 with 16K context) 24 GB32 GB DDR5 ~55 tok/s on 14B FP16 (RTX 4090) · ~78 tok/s (RTX 5090)
msi Gaming GeForce RTX 4090 24GB GDRR6X 384-Bit HDMI/DP… Tracked at $2,950 on 2026-09-16 — price may vary. View Current Price →

Phi-4 mini fits on almost any modern card. Quality is solid for a 4B model — beats Llama 3.2 3B on most benchmarks.

Frequently asked questions

How much VRAM do you need to run a 70B model locally?

About 48 GB resident at Q4_K_M. Q4 weights run roughly 0.55 GB per billion parameters, so 70B is ~39 GB of weights, and the runtime plus a usable context window wants about 2 GB on top — with headroom for longer context that lands at 48 GB. In practice that is two 24 GB cards pooled or one 48 GB workstation card. A single 24 GB card has to offload to system RAM, at which point PCIe bandwidth sets the speed rather than the GPU.

Can you run a 70B model on 12 GB of VRAM?

Not resident. 12 GB comfortably holds the 14B class at Q4 with room for context, and that is the ceiling for a card that size. A 70B model on 12 GB means offloading the majority of the weights to system RAM, which drops throughput by roughly an order of magnitude and turns the figure into a memory-bandwidth measurement rather than a GPU one. If 70B is the requirement, the card is the wrong axis to economise on.

How do you work out the VRAM a model needs?

Multiply the parameter count in billions by about 0.55 GB for Q4_K_M weights, then add roughly 2 GB for the runtime and a usable context window. A 32B model is therefore about 17.6 GB of weights and ~20 GB resident, which is why 24 GB is the practical floor for that class. Longer context costs more: KV-cache grows with context length and with batch size, so a 128K-context session on the same model wants materially more than the same model at 8K.

Is a lower quantisation worth it to fit a bigger model?

Usually down to Q4, rarely below. Published comparisons put Q4_K_M within about 1–2% of FP16 on benchmark averages, which is below what most workloads notice. Q3 starts to show — a 4–6% drop and visible failures on edge cases. Dropping a quantisation level to squeeze a larger parameter count onto the same card is generally a better trade than offloading; dropping two to do it usually is not.

Is unified memory an alternative to a discrete GPU for local LLMs?

For capacity, yes; for throughput, no. A Mac Studio with 192 GB of unified memory or an AMD Ryzen AI Max system loads model sizes no consumer NVIDIA card can hold, silently and at low power. Per-token speed is lower than a discrete card of equivalent tier because memory bandwidth, not capacity, sets throughput once the weights fit. The trade is capacity and quiet against tokens per second — see the per-model tiers below for where each lands.

Where to go next