RTX 3090 for Local LLMs: What It Runs, What It Costs, Every Guide
The RTX 3090's 24 GB holds models up to about DeepSeek-R1-Distill-Qwen 32B entirely in VRAM: its 19.9 GB Q4_K_M file leaves 4.1 GB for context (per the SpecPicks Will-It-Run matrix). New, it is $1,730 this week, $72.08 per GB of VRAM, per the Amazon listings SpecPicks tracks (price may vary).
What fits in 24 GB
Each row is the model’s real Q4_K_M GGUF file size against the card’s 24 GB, with the median of published, source-linked generation-speed runs that pass a memory-bandwidth sanity check. Open a row for the sources behind it.
| Model | Q4 file | Verdict | Median speed |
|---|---|---|---|
| Llama 3.2 1B | 0.8 GB Q4_K_M | Runs | 267 tok/s 1 cited run |
| Llama 3.2 3B | 2 GB Q4_K_M | Runs | no cited run |
| Mistral 7B | 4.4 GB Q4_K_M | Runs | no cited run |
| Qwen2.5 7B | 4.7 GB Q4_K_M | Runs | no cited run |
| Llama 3.1 8B | 4.9 GB Q4_K_M | Runs | 92 tok/s 3 cited runs |
| DeepSeek-R1-Distill-Llama 8B | 4.9 GB Q4_K_M | Runs | no cited run |
| Qwen3 8B | 5 GB Q4_K_M | Runs | 101 tok/s 2 cited runs |
| Gemma 2 9B | 5.8 GB Q4_K_M | Runs | no cited run |
| Gemma 3 12B | 7.3 GB Q4_K_M | Runs | no cited run |
| Qwen2.5 14B | 9 GB Q4_K_M | Runs | 55.8 tok/s 3 cited runs |
| DeepSeek-R1-Distill-Qwen 14B | 9 GB Q4_K_M | Runs | 37.5 tok/s 1 cited run |
| Qwen3 14B | 9 GB Q4_K_M | Runs | 61.1 tok/s 2 cited runs |
| Phi-4 14B | 9.1 GB Q4_K_M | Runs | no cited run |
| gpt-oss-20b | 12.1 GB MXFP4 | Runs | no cited run |
| Gemma 3 27B | 16.6 GB Q4_K_M | Runs | no cited run |
| Gemma 2 27B | 16.7 GB Q4_K_M | Runs | no cited run |
| Qwen3 30B-A3B | 18.6 GB Q4_K_M | Runs | no cited run |
| Qwen3 32B | 19.8 GB Q4_K_M | Runs | 32.7 tok/s 2 cited runs |
| Qwen2.5 32B | 19.9 GB Q4_K_M | Runs | 23 tok/s 2 cited runs |
| DeepSeek-R1-Distill-Qwen 32B | 19.9 GB Q4_K_M | Runs | no cited run |
| Llama 3.1 70B | 42.5 GB Q4_K_M | Offload | no cited run |
| Llama 3.3 70B | 42.5 GB Q4_K_M | Offload | no cited run |
| DeepSeek-R1-Distill-Llama 70B | 42.5 GB Q4_K_M | Offload | no cited run |
What it costs this week
- New (Amazon): $1,730 · $72.08 per GB of VRAM
- Used (eBay median, 7 days): withheld — 0 listings seen, fewer than 8
View Current Price →DetailsUsed on eBay
*Tracked snapshot read 2026-09-28, not a live quote: price may vary. Its row in this week’s LLM GPU Price Index →
Every RTX 3090 guide, by question
19 SpecPicks guides cover this card. They are grouped by the question they answer, most-read first.
Compared with other cards and machines 6 guides
- Dual RTX 3090 vs RTX 5090: Gaming vs AI Training
- RTX 3090 vs RTX 4090 for LLM Inference: Same 24GB (2026)
- Dual RTX 3060 12GB vs RTX 3090 24GB for Qwen2.5 32B (2026)
- RTX 3060 12GB vs RTX 3090 for Local LLMs (2026)
- RTX 3090 24GB vs RTX 4060 Ti 8GB in 2026: Used VRAM or New Efficiency?
- RTX A6000 48GB vs RTX 3090 24GB for Production LLM Inference (2026)
Backends, drivers and runtimes 3 guides
Specific models 6 guides
- How to run Qwen 3 32B on NVIDIA GeForce RTX 3090
- How to run Llama 3.1 70B on NVIDIA GeForce RTX 3090
- How to run DeepSeek-R1 32B on NVIDIA GeForce RTX 3090
- Qwen 3.6 27B at 2x tok/s: Luce DFlash on One RTX 3090
- How to run Qwen 3 14B on NVIDIA GeForce RTX 3090
- How to run Llama 3.1 8B on NVIDIA GeForce RTX 3090
Builds, buying and value 2 guides
Everything else 2 guides
Where to go next
- RTX 3090 vs Radeon RX 7900 XTX — the card it is most often cross-shopped against.
- RTX 3090 benchmarks — every gaming and AI run on file, with sources.
- Will it run? — any model against any card.
- AI rigs hub — complete local-AI builds.
As an Amazon Associate, SpecPicks earns from qualifying purchases.
This page is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.