As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices shown are catalog prices at the time of writing and may vary.
Why this comparison is about VRAM, not frame rates
The existing RTX 5090 vs RTX 4090 comparison on SpecPicks covers gaming. For local language models the question changes shape. Frame rates depend on shader throughput. Token generation mostly depends on how fast the card can stream model weights out of VRAM, and on whether those weights and the KV cache fit in VRAM at all.
That makes the decision mostly a question about models. With 24 GB, 8B–14B models fit with plenty of room, and 27B–32B models fit at 4-bit if you keep the context moderate. With 32 GB you can run a 32B model at 5- or 6-bit, or at 4-bit with a much longer context, before anything spills into system RAM. Neither card runs a 70B model comfortably on its own. This synthesis walks through the published measurements, the file-size math and the power budget so you can see which side of the 24 GB line your workload is on.
Key takeaways
- Bandwidth: 1,792 GB/s (5090) vs 1,008 GB/s (4090), per the GeForce 40 series and 50-series spec tables. Decode speed tracks this closely.
- Measured generation gap: on the llama.cpp CUDA scoreboard, Llama 2 7B Q4_0 generates at 300.40 tok/s on a 5090 vs 188.96 tok/s on a 4090 with flash attention on. That is about 59% faster.
- Prompt processing is nearly a tie at short context: 14,970 vs 14,771 tok/s pp512 on the same scoreboard.
- Capacity: Qwen3 32B Q8_0 is a 34.82 GB file (Qwen3-32B-GGUF). It fits neither card. Q5_K_M (23.21 GB) and Q6_K (26.88 GB) fit only the 5090 once you add any context.
- Power: 575 W total graphics power and a 1,000 W recommended PSU for the 5090 (NVIDIA), vs 450 W and 850 W for the 4090 (NVIDIA).
- Price: launch MSRPs were $1,999 and $1,599. As of September 2026 the 5090 sells far above MSRP, and even used 4090s trade around $2,500 (getpcparts).
Step 0 — what is the largest model you run every day?
Answer this before you compare anything else.
- 8B–14B models (Qwen3 8B/14B, Llama 3.1 8B, Gemma 3 12B). Both cards hold these at Q8 with long context. The 5090 is faster, but a 4090 is already well past reading speed. Hardware Corner has the 4090 at 84.4 tok/s on Qwen3 14B at 4K, so a 4090, or even a cheaper card, is enough here.
- 20B–32B models (gpt-oss 20B, Gemma 3 27B, Qwen3 32B). This is where the choice matters. At Q4 both cards work, but the 4090 runs out of room for long context. Hardware Corner's 4090 table has no Qwen3 32B result beyond 16K; the 5090 table goes to 32K.
- 70B models (Llama 3.3 70B). Neither card is the right tool. Llama 3.3 70B Q4_K_M is 42.52 GB (bartowski's GGUF repo), well over 32 GB. Look at the dual-GPU and 48 GB options in the dual RTX 3060 vs single GPU for Llama 70B guide.
If you answered 1, you're shopping in the wrong price bracket. If you answered 2, read on. If you answered 3, a single consumer card won't do it.
Spec-delta table
| Spec | RTX 4090 | RTX 5090 | Delta | LLM impact |
|---|---|---|---|---|
| VRAM | 24 GB GDDR6X | 32 GB GDDR7 | +8 GB (+33%) | Decides which quant and context fit |
| Memory bus | 384-bit | 512-bit | +33% | Part of the bandwidth gain |
| Memory bandwidth | 1,008 GB/s | 1,792 GB/s | +78% | Sets the decode (tok/s) ceiling |
| CUDA cores | 16,384 | 21,760 | +33% | Helps prompt processing and batching |
| Total graphics power | 450 W | 575 W | +125 W | PSU, cooling and electricity cost |
| Recommended system PSU | 850 W | 1,000 W | +150 W | May force a PSU upgrade |
| Launch MSRP | $1,599 | $1,999 | +$400 | Street prices in 2026 are far higher |
Sources: NVIDIA's RTX 4090 and RTX 5090 spec pages for cores, memory, power and PSU; Wikipedia's 40- and 50-series tables for bandwidth and MSRP.
The bandwidth gap is bigger than the VRAM gap. That matters because a single user's chat session is decode-bound. Each generated token needs one pass over the active weights, so tokens per second is roughly bandwidth divided by the bytes read per token. The 5090 should decode about 1.5–1.8× faster on any model that fits both cards, and the published numbers below land in that range.
Which models fit on 24 GB vs 32 GB?
File sizes below come from the Hugging Face tree API for each GGUF repo (Qwen3-32B-GGUF, Gemma 3 27B GGUF, Llama 3.3 70B GGUF, gpt-oss-20b GGUF). Fit ratings are SpecPicks arithmetic: the file plus roughly 1.5 GB of runtime buffers plus the KV cache for a 4K–8K context.
| Model / quant | File size | Fits 24 GB (4090)? | Fits 32 GB (5090)? |
|---|---|---|---|
| gpt-oss 20B MXFP4 | 12.11 GB | Yes, long context | Yes, long context |
| Gemma 3 27B Q4_K_M | 16.55 GB | Yes | Yes |
| Gemma 3 27B Q6_K | 22.17 GB | Tight, short context only | Yes |
| Gemma 3 27B Q8_0 | 28.71 GB | No | Yes, short context |
| Qwen3 32B Q4_K_M | 19.76 GB | Yes, up to ~16K | Yes, 32K+ |
| Qwen3 32B Q5_K_M | 23.21 GB | No (no room for KV) | Yes |
| Qwen3 32B Q6_K | 26.88 GB | No | Yes, moderate context |
| Qwen3 32B Q8_0 | 34.82 GB | No | No |
| Llama 3.3 70B IQ2_XS | 21.14 GB | Barely, heavy quality loss | Yes, heavy quality loss |
| Llama 3.3 70B Q2_K | 26.38 GB | No | Short context only |
| Llama 3.3 70B Q3_K_M | 34.27 GB | No | No |
| Llama 3.3 70B Q4_K_M | 42.52 GB | No | No |
The practical upshot: the 5090's extra 8 GB buys one quantization step on 27B–32B models (Q4 → Q5/Q6), or roughly double the context at Q4. It does not turn a 70B model into a comfortable single-card workload. A 2-bit 70B quant fits the 5090, but a quant that aggressive loses a lot of quality. Many users get better answers from a 32B model at Q6.
How much faster is the 5090 in tokens per second?
Two independent sources measure both cards on the same harness.
llama.cpp CUDA scoreboard (Llama 2 7B Q4_0, from discussion #15013):
| Card | Flash attention | Prompt (pp512) tok/s | Generation (tg128) tok/s |
|---|---|---|---|
| RTX 4090 | off | 11,992.70 | 186.21 |
| RTX 5090 | off | 14,073.41 | 290.02 |
| RTX 4090 | on | 14,770.63 | 188.96 |
| RTX 5090 | on | 14,970.15 | 300.40 |
Hardware Corner (Q4_K, March 2026 update; RTX 4090 page, RTX 5090 page), token generation in tok/s:
| Model | Context | RTX 4090 | RTX 5090 | 5090 advantage |
|---|---|---|---|---|
| Qwen3 8B | 4K | 141.3 | 200.4 | +42% |
| Qwen3 8B | 32K | 82.3 | 129.8 | +58% |
| Qwen3 8B | 128K | 33.8 | 58.8 | +74% |
| Qwen3 14B | 4K | 84.4 | 123.8 | +47% |
| Qwen3 14B | 32K | 55.4 | 82.4 | +49% |
| Qwen3 32B | 4K | 39.6 | 61.4 | +55% |
| Qwen3 32B | 16K | 34.4 | 50.9 | +48% |
| Qwen3 32B | 32K | — | 43.8 | 4090 has no result |
| gpt-oss 20B | 4K | 190.6 | 298.2 | +56% |
The two sources agree: on a model that fits both cards, the 5090 generates about 1.4–1.6× faster, and the gap widens as the context grows. Puget Systems saw a smaller gap in its launch review. It reported the 5090 leading the 4090 by "about 29%" in llama.cpp token generation on Phi-3 Mini, and flagged a suspected early-driver issue in prompt processing (Puget Systems). That test ran in February 2025. The 2026 scoreboard numbers above reflect newer drivers and llama.cpp builds.
One caveat on crowd-sourced leaderboards: single user submissions can be misleading. LocalScore currently lists one 5090 run below two 4090 runs on Qwen2.5 14B. That result is outside every other published comparison, so this synthesis doesn't rely on it.
Prefill vs generation: why the gap isn't uniform
LLM inference has two phases, and they stress different parts of the card.
- Prefill (prompt processing) runs the whole prompt through the model in large matrix multiplications. It is compute-bound. The 5090 has 33% more CUDA cores, but at short prompts both cards are so fast that the difference barely shows. On the scoreboard's pp512 test with flash attention on, they are within 2% of each other. At longer prompts the 5090 pulls ahead: Hardware Corner has Qwen3 8B prefill at 6,034 vs 3,560 tok/s at 32K.
- Generation (decode) produces one token at a time and re-reads the weights every step. It is bandwidth-bound, and this is where the 78% bandwidth advantage shows up as a 40–60% speed advantage.
For chat, generation speed dominates what you feel. For RAG pipelines and coding agents that feed in 20K-token prompts, prefill speed decides how long you wait for the first token. The 5090 is better at both, but the prefill advantage only matters once your prompts are long.
Context length and the KV cache
The KV cache is the memory that holds attention state for every token in the context. It grows linearly with context length, and it's what pushes a 32B model off a 24 GB card.
For Qwen3 32B, the published config.json specifies 64 layers, 8 key-value heads and a head dimension of 128. At FP16 that works out to 2 × 64 × 8 × 128 × 2 bytes = 256 KiB per token (SpecPicks arithmetic):
| Context | KV cache (FP16) | Qwen3 32B Q4_K_M total (19.76 GB + KV) | 4090 (24 GB) | 5090 (32 GB) |
|---|---|---|---|---|
| 4K | ~1.1 GB | ~20.9 GB | Fits | Fits |
| 16K | ~4.3 GB | ~24.1 GB | At the limit; needs KV quantization | Fits |
| 32K | ~8.6 GB | ~28.4 GB | Does not fit | Fits |
| 64K | ~17.2 GB | ~37.0 GB | No | Needs Q8 KV cache |
That table explains why Hardware Corner's 4090 page stops at 16K for Qwen3 32B while the 5090 page continues to 32K. You can stretch a 4090 by quantizing the KV cache to Q8 (roughly halving it), but that is a workaround, not headroom. If your work is long-document RAG or agentic coding on a 32B model, the 5090's extra 8 GB is the deciding spec.
Is the 5090's 575 W worth it?
NVIDIA rates the RTX 5090 at 575 W total graphics power and recommends a 1,000 W system PSU. It ships with an adapter for four PCIe 8-pin cables (NVIDIA). The 4090 is rated at 450 W and 850 W (NVIDIA). Both use the 16-pin 12V-2x6 / 12VHPWR connector family, which has a documented history of melting when not fully seated (Wikipedia: 12V-2x6 power connector issue). Seat it fully and avoid tight bends near the plug.
Perf per watt (SpecPicks arithmetic, rated TGP): on the scoreboard's generation test the 5090 produces 300.40 / 575 = 0.52 tok/s per rated watt, and the 4090 produces 188.96 / 450 = 0.42. On Qwen3 32B at 4K the figures are 61.4 / 575 = 0.107 and 39.6 / 450 = 0.088. Decode rarely pulls full TGP, so real efficiency is better on both cards, but the 5090 comes out ahead per watt either way.
Perf per dollar depends on what you actually pay. At launch MSRPs, the 5090 cost 25% more for 40–60% more tokens per second, which was a clear win. In September 2026 the market is different. Used 4090s average $2,589 over 30 days on eBay sold listings, and used 5090s average $4,252 (getpcparts 4090, getpcparts 5090). At those prices the 5090 costs about 64% more for about 50% more speed, so the 4090 edges ahead on tok/s per dollar for models that fit it. The 5090 still wins on capacity, which no amount of price math fixes.
Software: Blackwell needs CUDA 12.8 or newer; that release "adds compiler support for … SM_120" (CUDA 12.8 release notes). Ollama lists the RTX 5090 under compute capability 12.0 and notes a 570+ driver requirement for some cards (Ollama GPU docs). If you move an existing inference container from a 4090 to a 5090, rebuild it against a current CUDA base image.
Which RTX 4090 and RTX 5090 cards to buy
For LLM work, board partner differences matter less than they do for gaming. Clocks barely affect bandwidth-bound decode. Choose on cooling noise, physical size and price.
- ASUS TUF Gaming GeForce RTX 5090 32GB — a 3.6-slot air cooler with a vapor chamber. It is the default 5090 pick if your case fits it. Catalog price at the time of writing is well above MSRP, so check the live price.
- ASUS ROG Astral GeForce RTX 5090 32GB — a 3.8-slot, four-fan flagship. It is quieter under long sustained loads, but it is bigger and costs more. Buy it only if acoustics during overnight batch jobs matter to you.
- MSI GeForce RTX 4090 SUPRIM Liquid X 24G — an AIO-cooled 4090. It is a good fit for small cases and for an always-on inference box, where a blower or triple-fan air cooler would heat-soak.
- GIGABYTE GeForce RTX 4090 Gaming OC 24GB — a conventional triple-fan 4090 with an anti-sag bracket in the box.
New 4090 stock is thin in 2026 because production has ended, so the used market is where most buyers will find one. Inspect the 16-pin connector for discoloration and ask about the card's mining or rendering history before you pay.
Don't need either? The RTX 3060 12GB baseline
If Step 0 put you in the 8B–14B bracket, neither flagship is good value. The MSI Gaming GeForce RTX 3060 12GB runs Llama 2 7B Q4_0 at 76.92 tok/s generation on the same llama.cpp scoreboard, which is faster than anyone reads. Its 12 GB holds 8B models at Q8 and 14B models at Q4. Used 3060s averaged $293 in September 2026 on getpcparts. For the full trade-off see RTX 3060 12GB vs RTX 4090 for Local LLMs.
Common pitfalls
- Buying the 5090 for 70B models. Llama 3.3 70B Q4_K_M is 42.52 GB. The 5090 only runs 70B at 2-bit or with CPU offload, and offload drops generation to single-digit tok/s.
- Ignoring the KV cache. A model that loads fine can fail at 16K context. Size VRAM for the file plus the context you actually use.
- Stale containers on Blackwell. Images built against CUDA 12.4 won't target SM_120. Symptoms range from load failures to silent fallback to slower kernels.
- Undersized PSUs. An 850 W unit that ran a 4090 may trip on 5090 transients. NVIDIA recommends 1,000 W.
- Trusting one leaderboard entry. Single crowd-sourced runs vary with drivers, power limits and background load. Compare several sources that use the same harness.
Verdict matrix
Get the RTX 5090 if…
- You run 27B–32B models daily and want Q5/Q6 quality or 32K+ context.
- You're building a long-context RAG or coding-agent box where prefill at 16K–32K matters.
- Your PSU is already 1,000 W+ and your case fits a 3.6-slot card.
Get the RTX 4090 if…
- Your largest daily model is a 32B at Q4 with ≤16K context, or anything smaller.
- You can find a clean used card near $2,500 and want the better tok/s-per-dollar on models that fit.
- You want to stay under an 850 W PSU.
Get an RTX 3060 12GB (or another 12–16 GB card) if…
- You run 8B–14B models. A flagship gives you speed you can't read fast enough to use.
Recommended pick
For a buyer building a dedicated local-LLM machine in late 2026, the RTX 5090 is the better card. It generates 40–60% faster on every model both cards can run, per the llama.cpp scoreboard and Hardware Corner, and its 32 GB is the difference between running Qwen3 32B at a comfortable context and fighting the 24 GB wall. The counter-case is price. If a used RTX 4090 costs you roughly $1,700 less and your models fit in 24 GB with room for context, the 4090 is the rational buy. You give up speed, not capability.
Bottom line
The RTX 5090 beats the RTX 4090 for local LLMs on bandwidth (1,792 vs 1,008 GB/s), VRAM (32 vs 24 GB) and measured generation speed (about 1.5× on Qwen3 32B). The 4090 remains a strong 24 GB card for anything up to a 4-bit 32B model at moderate context. Neither is a 70B machine.
Related guides
- RTX 5090 vs RTX 4090 (gaming comparison)
- RTX 3090 vs RTX 4090 for LLM Inference: Same 24GB
- RTX 5070 Ti vs RTX 5090 for Local LLMs: 16GB vs 32GB
- RTX 3060 12GB vs RTX 4090 for Local LLMs
- RTX 5090 benchmarks · RTX 4090 benchmarks
Live price comparison
See the current prices side by side on the MSI RTX 4090 vs ASUS TUF RTX 5090 comparison, or go straight to the ASUS TUF RTX 5090 and MSI RTX 4090 SUPRIM Liquid X product pages. Prices change daily.
Citations and sources
- NVIDIA, GeForce RTX 5090 specifications — CUDA cores, memory, 575 W TGP, 1,000 W PSU. Accessed September 24, 2026.
- NVIDIA, GeForce RTX 4090 specifications — CUDA cores, memory, 450 W TGP, 850 W PSU. Accessed September 24, 2026.
- Wikipedia, GeForce RTX 50 series and GeForce 40 series — memory bandwidth and launch MSRPs. Accessed September 24, 2026.
- ggml-org, llama.cpp CUDA performance scoreboard, discussion #15013 — Llama 2 7B Q4_0 pp512/tg128 rows. Accessed September 24, 2026.
- Hardware Corner, RTX 4090 LLM benchmarks and RTX 5090 LLM benchmarks — Qwen3 and gpt-oss generation and prefill by context length. Accessed September 24, 2026.
- Puget Systems, NVIDIA GeForce RTX 5090 & 5080 AI Review — launch-era llama.cpp comparison. Accessed September 24, 2026.
- Hugging Face GGUF repositories: Qwen3-32B-GGUF, Qwen3-32B config, Gemma 3 27B GGUF, Llama 3.3 70B GGUF, gpt-oss-20b GGUF — file sizes and architecture. Accessed September 24, 2026.
- NVIDIA, CUDA 12.8 release notes and Ollama, GPU support — Blackwell software support. Accessed September 24, 2026.
- getpcparts, RTX 4090 used prices and RTX 5090 used prices — eBay sold-listing averages as of September 19, 2026. Accessed September 24, 2026.
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported. VRAM-fit, KV-cache and per-watt figures are SpecPicks arithmetic from the cited specifications.
