Hand-picked GPUs, pre-built workstations, and custom rigs for running Llama 3.1, Qwen, and DeepSeek-R1 locally. VRAM tiers from 16 GB budget to 96 GB pro.
By Mike Perry, AI Hardware Editor — Updated 2026-09-06
159,933Products evaluated
14,953Benchmark scores
2,741Brands tracked
DailyUpdated
SpecPicks earns a commission from qualifying Amazon purchases at no extra cost to you. How we pick →
What can I run? Model size to hardware, at Q4
Start from the model you want to run, not the card you want to buy. At Q4_K_M quantization
the weights take roughly 0.55 GB per billion parameters, and the runtime plus a usable
context window wants about 2 GB on top. Once the weights fit, generation speed is a
bandwidth question; until they fit, it is a "how much of this is executing on your CPU"
question, and the answer is always "too much".
Model size
Weights at Q4
VRAM you need
Cheapest card that fits
Fastest measured
3B (Llama 3.2 3B, Qwen 3 4B)Runs on almost anything with a discrete GPU, and usably on modern integrated graphics.
The VRAM column is derived from the quantization math above. Every tokens-per-second figure
is a median of community-reported runs held in the SpecPicks benchmark
database, restricted to cards that hold the weights and to chipsets with at least three
independent runs on file. Each card name links to its benchmark page with the per-run sources.
Read the full local-LLM GPU buying guide →How we source benchmark numbers
Can this GPU run this model? The Q4 compatibility grid
Read down your card, across to your model. Fits at Q4 means the weights and a
usable context window both live in VRAM. Tight fit means the weights fit and
the context window does not, so long prompts spill to system RAM.
CPU offload means the weights do not fit at all, and generation speed becomes a
question about your CPU rather than your GPU — which is why those cells carry no number.
Verdicts are derived from the quantization arithmetic stated in each row: Q4_K_M weights run
about 0.55 GB per billion parameters and the runtime plus a usable KV cache wants roughly
2 GB more. Every tokens-per-second figure is a median of
116 community-reported runs held in the SpecPicks benchmark
database, with at least 3 independent runs behind any number shown; card
names link to the per-run sources. Bands with no measured Q4 runs on these cards are left out of the grid rather than
filled with empty cells: 3B, 70B+.
The model-size table above covers every band, including the
ones only a 48 GB card answers.
Read the full local-LLM GPU buying guide →How we source benchmark numbers
Which GPU do you need to run a local LLM?
The short answer is VRAM, then bandwidth, then everything else. An 8B model at Q4 needs
about 6 GB of weights; a 27B-32B model at Q4 needs about 20 GB. Below the line where the
weights fit, tokens per second falls off a cliff no matter how fast the chip is, because
the layers that do not fit are executing on the CPU.
Generation speeds below are medians of 346 community-reported
benchmark runs held in the SpecPicks benchmark database (llama.cpp scoreboards, vendor
blogs, LocalLLaMA threads), not a single best case. Each GPU name links to its full
benchmark page with the per-run sources.
Under $300
Entry tier. 8-12 GB holds an 8B model at Q4 with room for context; anything larger runs on the CPU.
*Price sourced from Amazon.com. Last updated 2026-09-06. Price and availability subject to change.
⚔️ The Biggest Decisions for a 2026 AI Rig
RTX 5090 vs 4090 for raw tok/s. Mac Studio M3 Ultra vs RTX 5090 for capacity. Threadripper Pro vs Mac Studio for fine-tuning. The matchups every local-LLM builder is researching — with real benchmark data and a side-by-side spec comparison.
Best for: RTX 5090 — 32 GB VRAM for local LLM inference
Picked for rtx 5090 — 32 gb vram for local llm inference. It has solid buyer feedback. Ranked first in the Best AI Flagship GPU bracket (32GB-class) by the SpecPicks scoring algorithm (rating × log-of-review-volume, with category and price-band filters applied) — open the comparison table on the product page for side-by-side specs and the live Amazon listing for current price.
Best for: RTX 3090 + dual-card LLM rigs (under $900)
For buyers optimizing for rtx 3090 + dual-card llm rigs (under $900), this is the Best 24GB Used / Sweet-Spot GPU bracket leader. a strong rating, and a stable supply line on Amazon Prime make it the safest bet in the slot. Compare against the runners-up via the Compare tool before clicking through.
Best Budget LLM GPU goes to this product for buyers who match 12-16 gb vram — 7b comfortable, 13b quantized. ASRock holds a strong rating. Worth flagging: it trades against the higher-tier picks on raw performance but wins on price-to-feature ratio, which is why it stays in this slot through the year as prices on the flagships bounce around.
The Best Big NVMe for Model Storage nomination, from Sandisk. Best fit for 4-8 tb gen4/5 — llama, sdxl, dataset cache. Sandisk sits in the top quintile of Amazon ratings. The runner-up here is closer on paper than buyers usually expect — open the spec sheet on the product page before assuming this is the obvious choice.
Check current price on AmazonBest for: 1500W+ ATX 3.1 for dual-card builds
Best PSU for Multi-GPU: a strong default for 1500w+ atx 3.1 for dual-card builds. Corsair has solid buyer feedback. Corsair has shipped consistent revisions over the last 12 months without breaking-change drivers or firmware regressions, which is unusual at this price point — part of why it stays on this list.
Picked for open-frame, vertical, dual-gpu. LINKUP has solid buyer feedback. Ranked first in the Best PCIe Riser / Multi-GPU Mount bracket (PCIe 5.0-class) by the SpecPicks scoring algorithm (rating × log-of-review-volume, with category and price-band filters applied) — open the comparison table on the product page for side-by-side specs and the live Amazon listing for current price.
Llama 3.1 70B at q4_K_M quantization needs ~42 GB of VRAM. The RTX 5090 (32 GB) fits it with CPU offload at ~34 tok/s. For native inference: dual RTX 4090/5090, or Apple M3 Ultra with 128 GB+ unified memory.
How much VRAM for a home AI rig?
Starter rigs with 12-16 GB VRAM (RTX 4060 Ti 16GB, Arc B580) run 7B-14B models. Enthusiast 24 GB cards (RTX 4090) handle 32B natively. Pro tier needs 32 GB+ (RTX 5090) or Apple unified memory for native 70B. Workstation tier (405B, fine-tuning) needs 64 GB+ VRAM or 128 GB+ unified memory.
Is a Mac Studio M3 Ultra better than an RTX 5090 for AI?
Depends on the workload. M3 Ultra with 512 GB unified memory is the only consumer-tier option that holds 405B models in memory; it wins on memory capacity, silence, and power draw. RTX 5090 wins on raw tokens/sec for models that fit in 32 GB VRAM (5090 ≈ 34 tok/s on 70B q4 vs M4 Max at ≈12 tok/s). Pick M-series for capacity, NVIDIA for speed.
Can I run local LLMs on a gaming PC?
Yes — any modern gaming PC with 16 GB+ VRAM runs 7B-14B models well. A single RTX 4070 Ti Super (16 GB) handles Llama 3.1 8B at ~50 tok/s via Ollama, plenty for chat, coding assistants, and RAG. Beyond 32B you need dedicated AI hardware.
Should I buy a used RTX 3090 for AI?
Yes, if priced under ~$650. The 3090 has 24 GB VRAM (same as 4090), native NVLink for dual-card VRAM pooling, and is the community favorite for dual-GPU local-LLM builds. Check for fan/VRAM-temp issues before buying; ex-mining cards with rebuilt fans are fine if temps look clean.
Can I run Llama 3 8B on 8 GB of VRAM?
Yes, at Q4 quantization. An 8B model at Q4_K_M is about 5 GB of weights, which leaves roughly 3 GB on an 8 GB card for the KV cache — enough for a few thousand tokens of context. The 8 GB cards in the SpecPicks benchmark database land around 42-65 tok/s on 8B Q4. Past about 8K context you will start spilling into system RAM; a 12 GB card removes that ceiling.
Is the RTX 3060 12GB good for local LLMs?
It is the cheapest card that comfortably runs a 12-14B model. The 12 GB of VRAM is the point: it holds an 8B model at Q4 with a large context window, or a 13-14B model at Q4 with a normal one, where the faster 8 GB cards cannot. It is not fast — community runs put it near 29 tok/s on 12-14B Q4 against 58 tok/s for a 16 GB RTX 5070 Ti — but capacity decides what you can run and speed only decides how long you wait.
What is the minimum GPU for running a 70B model?
About 42 GB of VRAM at Q4, so a single 48 GB card (RTX A6000, RTX 6000 Ada, Radeon Pro W7900) or two 24 GB cards. Benchmark runs on file show roughly 11-18 tok/s on 48 GB cards and about 25 tok/s on an 80 GB H100. A 32 GB RTX 5090 runs 70B only with layers offloaded to system RAM, which drops it to around 18 tok/s; a 24 GB card offloading falls to about 8 tok/s.
What is the best GPU under $400 for local AI in 2026?
The Arc B580, at about $310 across the 2 listings SpecPicks tracks for it. Its 12 GB of VRAM is the reason: that holds models up to about 14B at Q4 with a usable context window; above that band the weights spill to system RAM and PCIe bandwidth sets the speed, not the GPU. It posts a median 40 tok/s on an 8B model at Q4 across 19 community-reported runs on file, against 61.6 tok/s and 8 GB for the NVIDIA GeForce RTX 3070 at about $385. Read the VRAM figure before the tokens-per-second figure — capacity decides which models run at all, speed only decides how long you wait.
What is the cheapest GPU that runs a 30B model entirely in VRAM?
The NVIDIA GeForce RTX 3090, at about $1,550 across 12 tracked listings. A 30-35B model at Q4_K_M is roughly 19 GB of weights, so 24 GB is the first tier that holds one with a usable context window; this card posts a median 29.2 tok/s in that band across 6 community-reported runs on file. Below 24 GB a 30B model runs only with layers resident in system RAM, where PCIe and DDR bandwidth set the speed rather than the GPU.
How We Pick
SpecPicks recommendations combine manufacturer spec data, aggregated benchmark results from public review sources (TechPowerUp, PassMark, Tom's Hardware, Geekbench, Phoronix, the LocalLLaMA community), live Amazon review feedback (ratings × review volume), and editorial judgment on price-to-performance. We update picks continuously as new silicon ships and prices move. Full methodology →