RTX 4000 SFF Ada Generation — Benchmarks & Specs
*Price sourced from Amazon.com. Price and availability subject to change.
Bottom line: how fast is the RTX 4000 SFF Ada Generation?
At 4K (High), the RTX 4000 SFF Ada Generation averages 144 fps in Apex Legends, per CpuTronic. For local LLM inference it generates 768.0 tokens/sec running qwen3:30b at Q4 under llama.cpp, per LocalLLaMA. In Geekbench 5 OpenCL it scores 124,926 points, per technical.city.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The RTX 4000 SFF Ada Generation is a graphics card from the Ada Lovelace family from NVIDIA. Data on this page draws on 1+ Amazon listings, 10 synthetic benchmark results, 12 community AI inference reports, 2 measured game frame-rate results, aggregated from public benchmark databases (TechPowerUp, PassMark, Geekbench, Cinebench) and the LocalLLaMA community. Read this page when shopping the RTX 4000 SFF Ada Generation, comparing it against other graphics cards in your build, or sizing it for a specific workload (gaming at 1080p/1440p/4K, productivity benchmarks, or local LLM inference).
Gaming Performance (measured FPS)
Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.
| Game | Resolution | Settings | Relative | Avg FPS | 1% low | Source |
|---|---|---|---|---|---|---|
| Apex Legends | 4K | High | 144 fps | — | CpuTronic 2025-04-01 | |
| Cyberpunk 2077: Phantom Liberty | 4K | RT Ultra RT on DLSS Quality | 68 fps | — | CpuTronic 2025-04-01 |
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| qwen3:30b | Q4 llama.cpp | 768.0 tok/s | — | LocalLLaMA 2026-02-25 | |
| qwen3:235b | Q4 ollama | 324.0 tok/s | — | LocalLLaMA 2026-02-27 | |
| llama3.2:1b | q4_K_M ollama | 189.0 tok/s | — | LocalScore AI 2024-10-01 | |
| llama3.2:1b | q4_K_M ollama | 189.0 tok/s | — | LocalScore (Mozilla Builders) 2025-04-14 | |
| llama3.2:1b | q4_K_M llama.cpp | 189.0 tok/s | — | LocalScore 2024-06-01 | |
| llama3.2:1b | q4_K_M llama.cpp | 189.0 tok/s | — | LocalScore.ai 2024-11-01 | |
| llama3.2:1b | q4_K_M llama.cpp | 189.0 tok/s | — | localscore.ai 2024-01-01 | |
| llama3.1:8b | q4_K_M llama.cpp | 58.6 tok/s | — | Hardware Corner 2024-09-29 | |
| llama3.1:8b | q4_K_M llama.cpp | 58.6 tok/s | — | Hardware Corner 2024-09-29 | |
| llama3.1:8b | q4_K_M llama.cpp | 58.6 tok/s | — | hardware-corner.net 2024-09-29 | |
| llama3.1:8b | q4_K_M ollama | 44.4 tok/s | — | LocalScore (Mozilla Builders) 2025-04-14 | |
| llama3.1:8b | q4_K_M llama.cpp | 44.4 tok/s | — | LocalScore.ai 2024-11-01 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| Geekbench 5 OpenCL | 124,926 points | technical.city 2023-06-01 | |
| Geekbench 5 Vulkan | 110,912 points | technical.city 2023-06-01 | |
| PassMark G3D Mark | 20,605 points | PassMark VideoCardBenchmark 2023-04-21 | |
| PassMark G3D Mark | 20,492 pts | PassMark 2026-04-20 | |
| PassMark G3D Mark | 20,492 points | PassMark Software 2023-04-21 | |
| 3DMark Time Spy | 13,990 points | 3DMark 2023-06-01 | |
| 3DMark Steel Nomad (DX12 Graphics Score) | 2,206 points | UL Benchmarks 2024-11-01 | |
| 3DMark Steel Nomad | 2,206 points | UL Benchmarks (3DMark) 2025-01-01 | |
| 3DMark Steel Nomad | 2,206 points | UL Benchmarks (Futuremark) 2024-01-01 | |
| PassMark G2D Mark | 1,081 pts | PassMark 2026-04-20 |
Products Featuring the RTX 4000 SFF Ada Generation
RTX 4000 SFF Ada Generation — Frequently Asked Questions
What is the RTX 4000 SFF Ada Generation best used for?
When was the RTX 4000 SFF Ada Generation released, and what was its launch MSRP?
Where do the benchmark numbers on this page come from?
Can the RTX 4000 SFF Ada Generation run local LLMs?
Where can I buy the RTX 4000 SFF Ada Generation?
Buying guides that rank the RTX 4000 SFF Ada Generation's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the RTX 4000 SFF Ada Generation
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.
- GTX 1660 VRAM: Why 8GB Doesn't Exist (Real Specs)
- One Monitor, Two PCs: Monitor KVM vs Dual-Input for a Gaming Rig and AI Box
- Arc B580 12GB vs RTX 3060 12GB: Which 12GB Card for 1440p in 2026?
- The Greatest GPU Awards: A Decade of Standout Cards
- Best Wired Gear for Tournament-Legal LAN Play in 2026
- Alien: Isolation on Steam Deck: Best Settings for 2025
- How to Filter Steam Deck's 20K+ Games by FPS, Price, HLTB
- Best Gear for Online D&D and Virtual Tabletop Nights in 2026
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- How to Build a Windows 98 Retro PC in 2026
- Best Retro Handhelds in 2026 — From $35 to $500
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- Best 1440p Gaming GPUs in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Best GPU for Llama 70B at Home in 2026: RTX 3060 12GB Stack vs Single Workstation Card
- RTX 5090 vs RTX 5080: Which Should You Buy in 2026?
- RX 9070 XT vs RTX 3060 12GB for Local LLM Inference (2026)
- Noctua NH-U12S vs ML240L RGB: Best Cooler for a 5800X
- AA-AgentPerf: What the New Agentic Inference Benchmark Means for Local Coding Rigs
- NVK Open-Source Vulkan Driver Gains Experimental DLSS on Linux
- Best AMD GPU for 4K Gaming in 2026
- CompactFlash as an IDE Boot Drive for Your Win98 Retro PC in 2026
- Motherboard Sales Fell 33% in 2026: Is DIY PC Building Dying?
- Qwen3.6-27B on Dual RTX 3060 12GB: The $400 30-50 tok/s Local LLM Build
- Best SATA, IDE & CompactFlash Adapters for Retro PC Data Recovery in 2026
- AMD Ryzen AI Halo ($4K) vs a DIY RTX 3060 Local-LLM Rig
- Imaging Your Big-Box CD-ROM Collection: A CompactFlash + IDE Workflow
- Best Controller for RetroPie & Pi Emulation: SN30 Pro vs DualSense
- RTX 3090 vs RTX 4090 for LLM Inference: Same 24GB (2026)
- Best GPU for Local Llama 3 8B Under $400: Why the RTX 3060 12GB Wins
- Best Local LLM You Can Run on 12GB of VRAM in 2026
- AutomationBench Cost Gap: What DeepSeek V4's 5-Cent Task Means for Local Agent Rigs
- Solar-Powered Raspberry Pi Bird Identifier Shows the Pi AI Camera's Reach
- What Non-Gaming Tasks People Actually Run on Steam Deck
- BOSGAME VTA-439: A Linux-Friendly Ryzen AI Mini PC
- Dual GPU Llama.cpp Speedup: What Actually Helps
- Best Wired Gaming Headset Under $50 for PC and Console in 2026
- llama.cpp vs vLLM for Single-User Local Chat in 2026: Which Wins on a 12GB GPU?
More buying guides from SpecPicks
Browse all buying guides →- Best Tools for Building and Repairing Retro PCs in 2026
- Best GPUs for 4K Gaming in 2026
- Best Gaming Mice for 2026
- Best Graphics Cards for Gaming in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best CPU Coolers for 2026
- Best AM5 Motherboards for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best CPUs for Gaming in 2026
- Best PC Cases for Building in 2026
- Best CPUs for Content Creators in 2026
- Best Controllers for PC Gaming in 2026
- Best NVMe External Enclosures for 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best GPUs for Running Local LLMs in 2026
- Best External SSDs for Content Creators in 2026
- Best Mechanical Keyboards for Gaming in 2026
- Best NVMe SSDs for Gaming in 2026
- Best Gaming Monitors for 2026
- Best 4K Monitors for Content Creators in 2026