RTX 5090 for Local LLMs: What It Runs, What It Costs, Every Guide
The RTX 5090's 32 GB holds models up to about DeepSeek-R1-Distill-Qwen 32B entirely in VRAM: its 19.9 GB Q4_K_M file leaves 12.1 GB for context, at a median 49.8 tok/s across 2 cited runs (per the SpecPicks Will-It-Run matrix).
What fits in 32 GB
Each row is the model’s real Q4_K_M GGUF file size against the card’s 32 GB, with the median of published, source-linked generation-speed runs that pass a memory-bandwidth sanity check. Open a row for the sources behind it.
| Model | Q4 file | Verdict | Median speed |
|---|---|---|---|
| Llama 3.2 1B | 0.8 GB Q4_K_M | Runs | no cited run |
| Llama 3.2 3B | 2 GB Q4_K_M | Runs | no cited run |
| Mistral 7B | 4.4 GB Q4_K_M | Runs | no cited run |
| Qwen2.5 7B | 4.7 GB Q4_K_M | Runs | no cited run |
| Llama 3.1 8B | 4.9 GB Q4_K_M | Runs | 150 tok/s 1 cited run |
| DeepSeek-R1-Distill-Llama 8B | 4.9 GB Q4_K_M | Runs | no cited run |
| Qwen3 8B | 5 GB Q4_K_M | Runs | 186 tok/s 1 cited run |
| Gemma 2 9B | 5.8 GB Q4_K_M | Runs | no cited run |
| Gemma 3 12B | 7.3 GB Q4_K_M | Runs | no cited run |
| Qwen2.5 14B | 9 GB Q4_K_M | Runs | 89.9 tok/s 1 cited run |
| DeepSeek-R1-Distill-Qwen 14B | 9 GB Q4_K_M | Runs | 89.1 tok/s 1 cited run |
| Qwen3 14B | 9 GB Q4_K_M | Runs | 124 tok/s 1 cited run |
| Phi-4 14B | 9.1 GB Q4_K_M | Runs | no cited run |
| gpt-oss-20b | 12.1 GB MXFP4 | Runs | 298 tok/s 1 cited run |
| Gemma 3 27B | 16.6 GB Q4_K_M | Runs | 47.3 tok/s 1 cited run |
| Gemma 2 27B | 16.7 GB Q4_K_M | Runs | no cited run |
| Qwen3 30B-A3B | 18.6 GB Q4_K_M | Runs | 234 tok/s 1 cited run |
| Qwen3 32B | 19.8 GB Q4_K_M | Runs | 58 tok/s 3 cited runs |
| Qwen2.5 32B | 19.9 GB Q4_K_M | Runs | 45.1 tok/s 1 cited run |
| DeepSeek-R1-Distill-Qwen 32B | 19.9 GB Q4_K_M | Runs | 49.8 tok/s 2 cited runs |
| Llama 3.1 70B | 42.5 GB Q4_K_M | Offload | no cited run |
| Llama 3.3 70B | 42.5 GB Q4_K_M | Offload | no cited run |
| DeepSeek-R1-Distill-Llama 70B | 42.5 GB Q4_K_M | Offload | no cited run |
What it costs this week
- New (Amazon): no quotable listing this week
- Used (eBay median, 7 days): withheld — 2 listings seen, fewer than 8
View Current Price →DetailsUsed on eBay
*Tracked snapshot, not a live quote: price may vary. Its row in this week’s LLM GPU Price Index →
Every RTX 5090 guide, by question
36 SpecPicks guides cover this card. They are grouped by the question they answer, most-read first.
Compared with other cards and machines 18 guides
- RTX 5090 vs RTX A6000 for Local LLMs: Speed (5090) or Capacity (A6000)?
- RTX 5090 vs RTX 5080: Which Should You Buy in 2026?
- RTX 5090 vs Mac Studio M4 Max for AI — which wins in 2026?
- Mac Studio M3 Ultra vs RTX 5090 for AI Inference in 2026
- RTX 5090 vs RTX 4090 — is the Blackwell upgrade worth it?
- RTX 5090 vs RTX 4090 at 1080p Competitive: Does the New Flagship Help Esports FPS?
- RTX 5090 vs RX 7900 XTX — does AMD still matter?
- Tenstorrent TT-QuietBox 2 (Blackhole) vs RTX 5090: Should LLM Builders Care?
- RTX 5090 Prebuilt vs a $700 RTX 3060 Local-LLM Box: What Extra VRAM Actually Buys
- RTX 5090 vs RTX 6000: Specs, Gaming, AI Compared
- RTX 5090 vs RX 9070 XT: Specs, Gaming & AI Compared
- RTX 5090 vs RTX 4080: Specs, Gaming & AI Compared
6 more
- RTX 5090 vs. M5 Max 128GB for Agentic Dev (2026)
- RTX 5090 vs RTX 3090: Specs, Gaming & AI Compared (2026)
- RTX 5090 vs RTX 5070: Specs, Value & Which to Buy
- Gunnir Arc B580 vs RTX 5090D on DeepSeek: The Budget AI-Rig Upset Explained
- RTX 5070 Ti vs RTX 5090 for Local LLMs: 16GB vs 32GB
- RTX 4090 vs RTX 5090 for Local LLM Inference: 24 GB vs 32 GB (2026)
Backends, drivers and runtimes 2 guides
Front-ends, agents and tools 2 guides
Specific models 5 guides
Builds, buying and value 2 guides
Everything else 7 guides
- RTX 5090 and Ryzen 7 9800X3D Benchmark Showdown: 2026 Performance Deep Dive
- RTX 5090 Benchmark Comparison: 4K Gaming, AI & Value (2026)
- RTX 5090 AIO Cooler: Liquid Cooling Options in 2026
- RTX 5090 AI Models: Performance, VRAM, Power
- RTX 5090 AI Cores: What the Specs Actually Mean in 2026
- RTX 5090 AI TOPS: What the Official Number Means
- Alienware Area-51 RTX 5090 Gets $2,580 Off — Worth It?
Where to go next
- RTX 5090 vs RTX 4090 — the card it is most often cross-shopped against.
- RTX 5090 benchmarks — every gaming and AI run on file, with sources.
- Will it run? — any model against any card.
- AI rigs hub — complete local-AI builds.
As an Amazon Associate, SpecPicks earns from qualifying purchases.
This page is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.