Llama 4 is the broadest-deployed open-weight family — every inference framework supports it day-one, and the 70B variant is the canonical "is this open model good enough for production" benchmark. The 17B mid-tier added in 4.1 makes single-GPU local hosting much more accessible than Llama 3.
Hardware tiers for running Llama 4
🧪
Minimum (8B at Q4)
8B Q4 fits in 8 GB but you want 12-16 GB for context room. 8B is a good lightweight assistant; coding and complex reasoning are limited.
*Price sourced from Amazon.com. Price and availability subject to change.
🏆
Recommended (17B / 70B at Q4)
RTX 5090 (32 GB) lets you run the new 17B variant fully and 70B with light CPU offload. RTX 4090 (24 GB) handles 17B comfortably; 70B needs more aggressive offload and slows to ~8 tok/s.
VRAM24 GB
System RAM64 GB DDR5
Throughput~45 tok/s on 17B (RTX 4090) · ~14 tok/s on 70B Q4 with offload (RTX 5090)
*Price sourced from Amazon.com. Price and availability subject to change.
⚡
Luxury (70B native, 405B with offload)
405B is too big for any consumer single-GPU rig. Mac Studio M3 Ultra (192 GB unified) is the only consumer machine that holds it in memory. Dual RTX 3090 NVLink runs 70B comfortably at 24+ tok/s.
VRAM48 GB
System RAM256 GB DDR5
Throughput~22 tok/s on 70B Q4 (dual RTX 3090) · ~3 tok/s on 405B Q4 (Mac M3 Ultra 192 GB)
Llama 4 added a 17B mid-tier (Llama 3 had 8B → 70B with nothing in between), expanded context to 256K, and added native vision support across all sizes. Inference performance per parameter is similar to 3.3.
Does Llama 4 405B run locally?
Only on Mac Studio M3 Ultra 192GB (or 96GB with aggressive offload). No consumer NVIDIA rig fits it natively. Dual RTX 3090 + 256 GB system RAM works at ~3 tok/s with heavy CPU offload — usable for batch jobs, not chat.
What's the cheapest "real" Llama 4 rig?
Used RTX 3090 ($800-1100) + 64GB DDR5 + Ryzen 7 7700X. Runs 17B Q4 at ~50 tok/s and 70B Q4 with offload at ~10 tok/s. Total build: $1500-1800.
More guides & deep dives from the SpecPicks archive