GPU Picker: the right graphics card for your VRAM, budget and workload
Pick a VRAM floor, a budget and what you are running. The picker ranks the GPUs SpecPicks holds published benchmark runs for and shows the median measured throughput behind every recommendation, with the run count and a source link in the row. Nothing is estimated from a capacity formula — a card without enough published runs to quote a median is left out.
Your picks
For local llm inference at 8 GB or more, the pick is the NVIDIA GeForce RTX 5090 (32 GB): a median of 186 tok/s on 7-9B (Llama 3.1 8B, Qwen 3 8B) at Q4, across 4 published runs from 3 sources, and a median of 130 FPS at 1440p on High/Ultra, across 11 published runs. It is tracked at about $1,999 (launch MSRP — no street listing currently prices within a sane band of it).
| # | Card | VRAM | Price | Local LLM | Where |
|---|---|---|---|---|---|
| 1 | NVIDIA GeForce RTX 5090 | 32 GB | $1,999 MSRP | 186 tok/s 7-9B (Llama 3.1 8B, Qwen 3 8B) · 4 runs / 3 src · Hardware Corner | Search Amazon → Benchmarks |
| 2 | AMD Instinct MI300X 192GB | 192 GB | $15,000 MSRP | 162 tok/s 7-9B (Llama 3.1 8B, Qwen 3 8B) · 5 runs / 4 src · KnightLi GPU Benchmark Scoreboard | Search Amazon → Benchmarks |
| 3 | NVIDIA H100 PCIe 80GB | 80 GB | $32,000 MSRP | 146 tok/s 7-9B (Llama 3.1 8B, Qwen 3 8B) · 5 runs / 4 src · XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) | Search Amazon → Benchmarks |
| 4 | NVIDIA A100 PCIe 80GB | 80 GB | $15,000 MSRP | 138 tok/s 7-9B (Llama 3.1 8B, Qwen 3 8B) · 7 runs / 5 src · Markaicode | Search Amazon → Benchmarks |
| 5 | NVIDIA RTX A5000 24GB | 24 GB | $1,999 MSRP | 136 tok/s 7-9B (Llama 3.1 8B, Qwen 3 8B) · 5 runs / 5 src · llama.cpp GitHub (CUDA scoreboard) | Search Amazon → Benchmarks |
As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices are the lowest currently tracked across the listings we follow for each chipset and may vary — check the listing before committing.
How to read this
Local-LLM throughput is a median over runs at Q4-class quantisation on one model band — the largest band your chosen VRAM floor actually holds, named in the table caption. Every card is quoted on that same band on purpose: an 8B model runs several times faster than a 27B one on the same silicon, so ranking each card on whichever band it happens to have runs for would put a thinly-measured card above a well-measured one. A card with no published runs on the reference band is left out rather than ranked on an easier one. Gaming figures are medians at High or Ultra with ray tracing excluded, quoted at 1440p where the card has coverage there and 1080p otherwise. Both are aggregations of other people’s published measurements, cited per row; see the methodology page for how runs are collected and filtered.
Want the full picture for one card? Every name in the table links to its benchmark page, which lists every run and the source it came from. For the model-size view of the same data — what fits in each card rather than which card is fastest — see the model-fit table on /ai-rigs and the best GPUs for local LLMs guide.
Frequently asked questions
- How does the GPU picker choose a card?
- It filters the GPU catalogue to cards with at least the VRAM you asked for and a tracked price inside your budget, then ranks what is left by median measured throughput — tokens per second at Q4 for local LLM use, average FPS at High/Ultra for gaming, and the average of both standings when you pick Both. Every LLM figure is quoted on the same model band (the largest one your VRAM floor holds), because an 8B model runs several times faster than a 27B one and ranking cards on different bands would compare numbers that are not comparable. A card with too few published runs on that band to quote a median is left out rather than estimated.
- Where do the tok/s and FPS numbers come from?
- They are medians over benchmark runs published by other people and aggregated here, with the run count, the source count and a link to one of the underlying sources shown in each row. SpecPicks does no first-party benchmarking and does not claim any.
- How much VRAM do I need for a local LLM?
- At Q4 quantisation a rough floor is 8 GB for an 8B model, 12 GB for a 12–14B model, 16 GB for a 20–27B model, 24 GB for a 30–35B model and 48 GB for a 70B model, each with a usable context window. Below the floor for a given model the weights spill to system RAM and PCIe bandwidth sets the speed rather than the GPU.
- Are the prices live?
- Prices are the lowest currently tracked across the Amazon and eBay listings SpecPicks follows for that chipset, filtered to listings that price within a sane band of the card’s MSRP. They move constantly — treat them as a guide and check the listing before buying. Where no listing qualifies, the launch MSRP is shown and labelled as such.
- Is the picker free, and do I need an account?
- It is free and needs no account. Every result has its own URL, so you can bookmark a set of answers or paste it into a thread and it will open the same way for anyone else.
Other ways in
Build a whole machine with the PC Builder, compare two cards side by side on the comparison pages, or browse buying guides and the AI rigs hub.