NVIDIA GeForce RTX 4090 — Benchmarks & Specs
*Price sourced from Amazon.com. Price and availability subject to change.
Bottom line: how fast is the NVIDIA GeForce RTX 4090?
At 1440p (Ultra), the NVIDIA GeForce RTX 4090 averages 194 fps in Shadow of the Tomb Raider, per HyperCyber. For local LLM inference it generates 2550.0 tokens/sec running llama3.1:8b at FP16 under vllm, per Spheron Blog. In Deep Learning Benchmark (ResNet-50) it scores 128,000 images/sec, per Puget Systems. Its 24 GB of VRAM is the binding constraint for local inference: that capacity fits 32B-parameter models at Q4 without offloading to system RAM.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA GeForce RTX 4090 is a graphics card from the Ada Lovelace family released in 2022 from NVIDIA. Key on-paper specs include 24 GB of GDDR6X VRAM, 450W TDP. It launched with a $1,599 MSRP, though street prices typically diverge meaningfully from launch pricing — see the linked product cards below for current Amazon listings. Data on this page draws on 8+ Amazon listings, 10 synthetic benchmark results, 12 community AI inference reports, 14 measured game frame-rate results, aggregated from public benchmark databases (TechPowerUp, PassMark, Geekbench, Cinebench) and the LocalLLaMA community. Read this page when shopping the NVIDIA GeForce RTX 4090, comparing it against other graphics cards in your build, or sizing it for a specific workload (gaming at 1080p/1440p/4K, productivity benchmarks, or local LLM inference).
Gaming Performance (measured FPS)
Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.
| Game | Resolution | Settings | Relative | Avg FPS | 1% low | Source |
|---|---|---|---|---|---|---|
| Counter-Strike 2 | 1080p | Ultra Native | 312 fps | 290 fps | Phoronix 2023-12-10 | |
| Shadow of the Tomb Raider | 1440p | Ultra | 194 fps | — | HyperCyber 2022-10-12 | |
| Cyberpunk 2077: Phantom Liberty | 1440p | Ultra Native | 192 fps | — | TechSpot 2023-09-26 | |
| Alan Wake 2 | 1440p | Path Traced RT on DLSS Quality | 170 fps | — | Tweaktown 2023-10-26 | |
| Alan Wake 2 | 1440p | RT Ultra + Path Tracing + DLSS 3.5 FG RT on DLSS | 170 fps | — | Tom's Hardware 2023-10-24 | |
| Alan Wake 2 | 1080p | Ultra (High, RT Off) Native | 154 fps | — | Hardware Times 2024-02-24 | |
| 007 First Light | 4K | Ultra DLSS 4.5 | 145 fps | 125 fps | PCBench 2025-06-01 | |
| Black Myth: Wukong | 1440p | Cinematic (High quality, RT Off) Native | 138 fps | — | TechSpot 2024-08-19 | |
| Alan Wake 2 | 4K | Path Traced RT on DLSS Quality | 134 fps | — | Tweaktown 2023-10-26 | |
| Alan Wake 2 | 4K | RT Ultra + Path Tracing + DLSS 3.5 FG RT on DLSS | 134 fps | — | Tom's Hardware 2023-10-24 | |
| Alan Wake 2 | 4K | Maximum, Full Ray Tracing RT on DLSS 3.5 | 134 fps | — | Tom's Hardware 2023-10-24 | |
| 007 First Light | 1080p | Ultra | 128 fps | 115 fps | PCBench 2025-06-01 | |
| Crimson Desert | 1080p | Ultra | 124 fps | 103 fps | PCBench 2025-06-01 | |
| 007 First Light | 1440p | Ultra | 111 fps | 91 fps | PCBench 2025-06-01 |
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| llama3.1:8b | FP16 vllm | 2550.0 tok/s | 18.0 GB | Spheron Blog 2026-05-03 | |
| qwen3:32b | AWQ vllm | 650.0 tok/s | 22.0 GB | Spheron Blog 2026-05-03 | |
| llama3:8b | — ollama | 440.0 tok/s | — | LocalLLaMA 2026-04-01 | |
| qwen3:30b-moe | Q4_K_XL llama.cpp | 195.8 tok/s | 16.5 GB | Hardware Corner 2025-11-06 | |
| Llama 2 7B | INT4 (AWQ) llama.cpp | 194.0 tok/s | — | arXiv (LLM Inference Hardware Survey) 2024-10-06 | |
| llama2:7b | q4_0 llama.cpp | 189.0 tok/s | — | llama.cpp GitHub Discussion #15013 2025-08-01 | |
| Llama 3.1 8B | Q4_K_M llama.cpp | 165.0 tok/s | 5.5 GB | llama.cpp GitHub 2024-08-22 | |
| Llama 3 8B | unspecified llama.cpp | 150.0 tok/s | — | NVIDIA Developer Blog 2024-10-01 | |
| qwen3:32b | q4_K_XL llama.cpp | 139.7 tok/s | — | Puget Systems 2025-06-01 | |
| llama3.1:8b | Q4_K_XL llama.cpp | 131.0 tok/s | 4.8 GB | Hardware Corner 2025-11-06 | |
| llama3.1:8b | q4_K_M llama.cpp | 125.0 tok/s | 5.5 GB | MyAIHardware 2025-05-01 | |
| llama3.1:8b | q4_K_M llama.cpp | 113.0 tok/s | 6.2 GB | Awesome Agents LLM Leaderboard 2025-06-01 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| Deep Learning Benchmark (ResNet-50) | 128,000 images/sec | Puget Systems 2024-12-10 | |
| 3DMark Steel Nomad | 46,384 points | The FPS Review 2024-05-20 | |
| 3DMark Time Spy Extreme | 38,450 pts | TechPowerUp 2024-11-02 | |
| PassMark G3D Mark | 38,066 pts | PassMark 2026-04-20 | |
| 3DMark Time Spy | 36,560 points | [H]ard|Forum RTX 4090 Time Spy Thread 2023-01-01 | |
| 3DMark Time Spy GPU | 35,200 pts | TechPowerUp 2023-11-15 | |
| 3DMark Port Royal | 27,901 points | HWBot 2023-01-01 | |
| 3DMark Port Royal | 24,886 points | Overclock3D 2022-10-12 | |
| 3DMark Time Spy Extreme | 20,192 points | The FPS Review 2022-09-09 | |
| 3DMark Time Spy Extreme (Graphics) | 19,000 pts | TechPowerUp 2022-09-01 |
Products Featuring the NVIDIA GeForce RTX 4090
Full Specifications
| tdp w | 450 |
|---|---|
| vram gb | 24 |
| vram type | GDDR6X |
| cuda cores | 16384 |
| boost clock mhz | 2520 |
NVIDIA GeForce RTX 4090 — Frequently Asked Questions
What is the NVIDIA GeForce RTX 4090 best used for?
When was the NVIDIA GeForce RTX 4090 released, and what was its launch MSRP?
Where do the benchmark numbers on this page come from?
Can the NVIDIA GeForce RTX 4090 run local LLMs?
Where can I buy the NVIDIA GeForce RTX 4090?
Buying guides that rank the NVIDIA GeForce RTX 4090's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA GeForce RTX 4090
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.
- RTX 3090 vs RTX 4090 for LLM Inference: Same 24GB (2026)
- RTX 3060 12GB Local LLM Guide: Which Models Actually Fit (2026)
- M3 Ultra vs RTX 4090 for local LLMs
- Mistral Medium 3.5 Local Inference: Hardware Requirements and Benchmarks
- PFlash on a Single RTX 3090: 10× Prefill Speedup at 128K Context vs llama.cpp
- How to run Llama 3.1 8B on NVIDIA GeForce RTX 4090
- How to run Qwen 3 14B on NVIDIA GeForce RTX 4090
- How to run DeepSeek-R1 32B on NVIDIA GeForce RTX 4090
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- How to Build a Windows 98 Retro PC in 2026
- Best 1440p Gaming GPUs in 2026
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- Best Retro Handhelds in 2026 — From $35 to $500
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Best 24GB GPU for Local LLM Inference in 2026
- Claude Sonnet 5: What Shipped and What It Means for Local Rigs
- Rescue Your Retro PC's Data Before the IDE or SATA Drive Dies
- 3dfx Voodoo2 SLI in 2026: Build Guide, Glide Setup, CF Storage
- Ryzen 5 5600G vs Ryzen 7 5700X for a Budget LLM + Gaming Build
- Qwen3.6 the Right Way: Run It Through a Pi Coding Agent
- The vLLM MCP Vulnerability: What Local LLM Operators Need to Do
- Best CompactFlash + IDE Adapters for Retro PC Builds in 2026
- Grok 4.3 vs GPT-5 vs Claude 4.7: Local Hardware Implications of the Closed-Model Intelligence Index
- Best SSD for a Raspberry Pi 4 Media Server in 2026
- Logitech G29 vs HORI vs Thrustmaster: Best Sim Racing Wheel for Beginners
- Build a Private Security Camera on a Raspberry Pi (Ring Alternative) in 2026
- Does Dual-Channel RAM Matter for Local LLM Inference?
- Jetson Orin Nano Super for Local LLM: 7B/13B Tokens-per-Second Reality Check
- Run a Local Coding Agent on an RTX 3060 12GB (After Codex Went Autonomous)
- Raspberry Pi 4/5 Wall Mount Case: 3D Printing Guide
- Gemma 4 31B Abliterated on a Single RTX 3060 12GB: Quantization, VRAM, and Real Tok/s
- RTX 4090 AIO Kit 2025: Pricing, Cooling & Buying Guide
- Bilingual Voice Agents: Frontier ASR on Code-Switched Speech
- Raspberry Pi Zero 2W: Privacy-Preserving Ring Alternative
- Linux Gaming Keeps Closing the Gap: Kernel and Windows-API Work Behind the Gains
- Transcend CF133 CompactFlash as a Win98 Boot Drive: Speed Tests, Compatibility, and the Adapter That…
- Intel Arc B580 & Arc Pro B60 24GB vs RTX 3060 12GB for Local LLMs
- Best Dock and Storage Combo for Handheld PC Gaming: JSAUX Dock + a Fast SSD
More buying guides from SpecPicks
Browse all buying guides →- Best Controllers for PC Gaming in 2026
- Best PC Cases for Building in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best Graphics Cards for Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best 4K Monitors for Content Creators in 2026
- Best CPU Coolers for 2026
- Best NVMe SSDs for Gaming in 2026
- Best Gaming Monitors for 2026
- Best GPUs for 4K Gaming in 2026
- Best CPUs for Content Creators in 2026
- Best Gaming Mice for 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best External SSDs for Content Creators in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best CPUs for Gaming in 2026
- Best AM5 Motherboards for 2026
- Best NVMe External Enclosures for 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best Mechanical Keyboards for Gaming in 2026