Skip to main content
RTX 5090 AI Build Guide: CPU, RAM, PSU & Cooling for Local Inference

RTX 5090 AI Build Guide: CPU, RAM, PSU & Cooling for Local Inference

CPU, memory, power and cooling picks for a 575W-class local-inference workstation.

The full parts list for an RTX 5090 AI build: CPU pairing, PSU headroom, DDR5 capacity and AIO cooling for local inference in 2026.

An RTX 5090 local-inference build wants a modern 8-16 core CPU, 64GB of DDR5, a 1000W-class ATX 3.1 power supply with a native 12V-2x6 connector, and either a large air cooler or a 240mm+ AIO on the CPU. Budget roughly $3,500 all-in with the card. The GPU is the star; the rest of the parts list exists to keep it fed without becoming the bottleneck or the failure point.

Editorial intro: who buys a 5090 for AI

The 5090 arrived in early 2025 with 32GB of GDDR7 and, more importantly for local-LLM work, memory bandwidth north of 1.7 TB/s per NVIDIA's 50-series page — roughly double the RTX 4090 it replaced. That bandwidth is what makes a Q4-quantized 32B model like Qwen 3.6 32B or Llama 3.5 34B run at 40-50 tok/s on a single card. It's also what makes the card 575W TGP and, therefore, an interesting build challenge.

The audience for a 5090-centric AI build is not the tinkerer with a 12GB card who wants to run 27B models — that person is served by our Qwen 3.6 on RTX 3060 guide and probably shouldn't be spending $2,000+ on a GPU. The 5090 build is for the developer or small-team ML engineer who wants a workstation that runs the current generation of 30B/70B-quantized models at productive speeds without spinning up cloud GPUs, and who's willing to spec the rest of the rig around one very hot, very power-hungry card. This guide walks that parts list, notes where the money is well spent, and — importantly — flags the places builders reflexively over-spec.

If you want an on-ramp with the same platform and lower up-front cost, an AMD Ryzen 7 5800X plus MSI RTX 3060 Ventus 3X 12G build shares the same CPU, PSU, and cooling class and lets you drop in the 5090 later. That's a legitimate strategy given 5090 supply, and we'll return to it in the bottom line.

Key takeaways

  • CPU class: 8-16 cores at high boost. A Ryzen 7 5800X, Ryzen 9 7900X, or Core i7-14700K are all fine — this workload does not reward a Threadripper.
  • RAM: 64GB DDR5-5600 minimum, 96GB or 128GB if you plan to CPU-offload 70B models. Dual-channel matters more than capacity beyond 64GB.
  • PSU: 1000W minimum with ATX 3.1 and a native 16-pin 12V-2x6 cable. Do NOT use an adapter.
  • Cooling: Air (Noctua NH-D15 class) or a good 240mm AIO like the Cooler Master ML240L on the CPU; the GPU cools itself. A 360mm is only warranted for very high-core CPUs sustained at 100%.
  • Storage: A single 1TB Samsung 970 EVO Plus NVMe or newer PCIe 4.0 drive is fine; larger models want 2TB to hold multiple quants comfortably.
  • Case: 45L+ ATX with good top exhaust. The 5090's 3-slot cooler is huge; measure before you buy.

Step 0 — diagnose the workload

The single biggest parts-list mistake with a 5090 AI build is speccing for a workload the builder doesn't actually have.

  • Inference-only chat/RAG: 8-core CPU, 64GB RAM, single NVMe. The GPU does everything. Overspending on CPU here is money set on fire.
  • Fine-tuning small (7B) models on a single 5090: Still fine on 8-16 cores, but you'll want 96-128GB of RAM to hold datasets in memory.
  • Fine-tuning larger (30B+) models with LoRA/QLoRA: You want fast NVMe for checkpoints and 128GB of RAM. Still no need for Threadripper.
  • Gaming + AI on the same rig: This is where high single-thread perf earns its keep. A 7800X3D or 14700K makes sense over a 5800X.
  • Multi-GPU expansion: Different story entirely. You need PCIe lane counts and a HEDT-class platform. Not the 5090-single-card guide.

If you don't need multi-GPU or heavy CPU-side data pipelines, the workload is bandwidth-bound to the GPU and everything else exists to keep that GPU fed.

Which CPU pairs sensibly with an RTX 5090 for inference?

Community measurements collected on r/LocalLLaMA show that single-user local inference is remarkably CPU-insensitive when the model fits fully in VRAM. Going from a 6-core Ryzen 5 5600G to a 16-core 9950X changes generation throughput by a few percent at most, per the same workload — the CPU's job is orchestration, not compute.

CPUCores/ThreadsApprox MSRPFit for 5090 inference
AMD Ryzen 5 5600G6/12$130Fine for pure inference; iGPU is a bonus
AMD Ryzen 7 5800X8/16$180-220Sweet spot; excellent single-thread + 8 fast cores
AMD Ryzen 7 7800X3D8/16$349Overkill for AI, ideal if you also game
AMD Ryzen 9 7900X12/24$349Good for data prep + inference combo
Intel Core i7-14700K20 (8P+12E)/28$349Solid all-rounder
AMD Ryzen 9 9950X16/32$549Only if you do heavy CPU-side ML work

The honest recommendation: buy the Ryzen 7 5800X if you're primarily doing inference. It's mature, cheap, and does not become the bottleneck. Save the CPU-budget delta for the 5090 you're actually optimizing around.

How much system RAM does a 5090-class build need?

The 5090 has 32GB of VRAM — enough to hold a Q4_K_M-quantized 34B model with 8k context fully resident. For that workload, system RAM is orchestration only, and 64GB is plenty.

The two scenarios that push you past 64GB:

  1. CPU-offloaded 70B models. A Q4_K_M 70B model weighs ~40GB — most of it lives in RAM if you want to load it at all. Add OS + tooling overhead and you're using 60-64GB of system RAM before you start.
  2. Heavy fine-tuning data pipelines. Batching 100k+ conversational examples in memory beats hitting SSD every epoch.
RAM configRecommended for
32GB DDR5Not enough — you'll swap under any real workload
64GB DDR5-5600Baseline 5090 inference build
96GB DDR5-5600Good balance for occasional 70B offload
128GB DDR5-5600Fine-tuning + oversized CPU-offloaded models

Skip the "buy fastest RAM you can" reflex. DDR5-5600 CL36 is well under $200 for a 64GB kit and 6000 CL30 costs 20-30% more for a real-world 2-3% inference gain. That money buys better cooling.

PSU wattage and connectors for a 575W-TGP card

The 5090's rated TGP is 575W with reported transient spikes to 900W+, according to reviews from Tom's Hardware. Add ~200W for a modern CPU + drives + fans and the rated system draw sits around 800W under a worst-case load.

PSU sizing rule: total sustained draw + 30-40% headroom for transients. That puts a 5090 build at 1000W minimum, with 1200W a safe pick for anyone who wants the transient margin. Older PSUs work electrically but produce two failure modes:

  1. Melting 12VHPWR connectors. The original 12VHPWR spec had genuine contact-quality problems and burned a handful of 4090 cables. The ATX 3.1 revision — 12V-2x6 — moves the sense pins and is safer. You want a PSU with a native 12V-2x6 cable, not an adapter.
  2. Transient shutdowns. Older PSUs OCP-trip on the 5090's spikes even when average draw is well below rating. Buy an ATX 3.1 unit rated for GPU transients.

Best-practice PSU spec for a 5090 build:

  • 1000-1200W
  • ATX 3.1 compliant
  • Native 12V-2x6 cable (no adapter)
  • 80 Plus Gold or better
  • Single-rail preferred for GPU-heavy builds

Do not skimp here. This is the component whose failure kills every other component.

Spec-delta table: CPU options for the build

Community-published inference measurements aren't always apples-to-apples, but the pattern is consistent: with the model fully in VRAM, generation throughput varies within 5% across sane modern desktop CPUs.

CPUBase/Boost GHzTDPApprox tok/s (Q4 34B on 5090)
Ryzen 5 5600G3.9/4.465W~48-50
Ryzen 7 5800X3.8/4.7105W~50-52
Ryzen 7 7800X3D4.2/5.0120W~51-53
Ryzen 9 7900X4.7/5.6170W~51-53
Core i7-14700K3.4/5.6125W~51-53

Those tok/s numbers are noisy — the takeaway is that on inference-only workloads the CPU is not the bottleneck. Prompt-eval throughput scales better with CPU cores but even there the delta is 10-15%, not 2x.

Benchmark table: expected tok/s tiers by model size on 32GB VRAM

Rough tiers for the 5090 based on community measurements of comparable 32-38B GGUF models on the card:

Model classQuantFits in VRAM?Gen tok/s (approx)
7B (Llama 3.5, Mistral 8B)Q8_0Yes100-130
14B (Qwen 3.6 14B)Q6_KYes65-85
27B (Qwen 3.6 27B)Q5_K_MYes45-55
34B (Llama 3.5 34B)Q4_K_MYes40-50
70B (Llama 3.5 70B)Q4_K_MNo (offload)6-12

A 5090's honest workload sweet spot is 14B-34B models at high quant. It runs 70B models via CPU offload, but the moment you spill to system RAM the throughput drops to what a 3090 does — cheaper hardware would have gotten you there.

Cooling: air vs 240mm AIO

The 5090 cools itself. Your cooling budget is for the CPU and case airflow.

Cooler classFit for the build
120mm air (Wraith Prism)Not enough for a 5800X sustained
Noctua NH-U12S / D15Air-cooling gold standard; handles 105-125W CPUs quietly
Cooler Master ML240LSolid 240mm AIO; matches D15 in a smaller footprint
360mm AIOOverkill for 5800X; sensible for 7950X/9950X sustained

If you're chasing quiet, the Cooler Master MasterLiquid ML240L puts the fans up top and out of the GPU's exhaust column, which noticeably lowers case temps when the 5090 is dumping 500W of heat sideways. The GPU is the room heater; keep the case's top and rear paths clear and don't stack the CPU cooler where it re-ingests the GPU's hot air.

Perf-per-dollar: 5090 build vs a dual-3060 budget alternative

If you priced this out today (mid-2026):

BuildGPU spendTotal system34B Q4 tok/s
Single RTX 5090~$2,000~$3,300~45
Single RTX 3090 (used)~$500~$1,700~20
Dual RTX 3060 12GB (24GB pooled)~$520~$1,700~15-18
Single RTX 4090 (used)~$1,500~$2,700~35

The 5090's headline win versus a used 3090 is roughly 2.2x the throughput for 4x the money. That's a real regression on pure perf-per-dollar. What it buys you: 32GB of VRAM (34B Q4 comfortably), current architecture (fp8, tensor-core generation 5), full warranty, and no need to source used cards.

Dual-3060 rigs are miserable to configure for LLM inference. VRAM doesn't pool the way GPU memory does across model-parallel workloads — you need either tensor parallelism (llama.cpp supports it partially, vLLM does better) or run separate model instances. It's a legitimate answer for hobbyists; it is not a productivity build.

Common pitfalls

  1. Buying a 750W PSU because the RTX 5090 "rated draw is 575W". Transient spikes trip OCP; average draw isn't the metric to size against.
  2. Adapting an old 12VHPWR cable. Even if it "fits", the sense-pin geometry changed. Buy an ATX 3.1 PSU with a native 12V-2x6 or accept a small but real fire risk.
  3. 32GB of RAM. Every 5090 build with 32GB ends up upgrading within a month once the owner tries a 70B offload.
  4. A 4-slot case. The 5090 reference cooler is 3 slots and 336mm long. Measure the case. Twice.
  5. Overspending on the CPU. A 9950X does not make inference faster than a 5800X. It makes you $370 poorer.
  6. DDR5-8000+ RAM kits. The XMP/EXPO stability at extreme speeds isn't worth the bragging rights on this workload.

When NOT to build a 5090-centric rig

  • You only run 7B/8B models. A 12GB 3060 or even an M-series Mac Mini is dramatically better $/tok.
  • You're renting cloud GPUs at $2/hour and running <1000 hours a year. Cloud math wins.
  • You want to fine-tune 30B+ models seriously. A 5090 handles small LoRA runs but you'll want an H100/A100 for real training.
  • You're on a stock 550W PSU and 350L case. The path to a 5090 build includes replacing those; that's a bigger project than swapping a card.

Bottom line: the balanced parts list

For a productivity local-inference workstation targeting 30B/34B-class models at Q4:

  • GPU: RTX 5090 32GB
  • CPU: AMD Ryzen 7 5800X or a 7800X3D if the rig also games
  • RAM: 64GB DDR5-5600 CL36 (upgrade to 96/128 if 70B is on the roadmap)
  • PSU: 1000-1200W ATX 3.1 with native 12V-2x6 (do NOT downgrade)
  • CPU Cooler: Cooler Master ML240L or Noctua NH-D15
  • Storage: Samsung 970 EVO Plus 1TB NVMe as scratch, plus a 2TB PCIe 4.0 drive for model weights
  • Case: 45L+ ATX with good top exhaust; a mesh front panel is doing real work with a 5090 inside
  • Backup path: If a 5090 is out of stock (still happening), buy the CPU/PSU/RAM/case now with a Ryzen 5 5600G or Ryzen 7 5800X and an MSI RTX 3060 Ventus 3X 12G and drop the 5090 in later.

Skip the $349 CPU and the $250 360mm AIO. Spend the delta on RAM headroom or the 2TB weight drive.

Related reading

Citations and sources

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Tap any product for full specs, live Amazon & eBay pricing, and alternatives.

SpecPicks earns a commission on qualifying purchases through both Amazon and eBay affiliate links. Prices and stock update independently.

Watch a review

Friendly Fire: AMD Ryzen 7 5800X CPU Review & Benchmarks vs. 5600X & 5900X — Gamers Nexus on YouTube

Frequently asked questions

Do I need the newest CPU for an RTX 5090 AI build?
For inference-dominated workloads the GPU does the heavy lifting, so a strong 8-core like the Ryzen 7 5800X keeps the card fed for single-user local LLM work without becoming the bottleneck. You'd only chase a newer high-core platform if you also run heavy CPU-offloaded models, large data-prep pipelines, or plan multi-GPU expansion where PCIe lanes and memory bandwidth start to matter more.
How big a PSU does a 575W-class GPU need?
Plan the supply around the card's rated TGP plus transient spikes and the rest of the system, which pushes most 5090-class single-GPU builds toward a high-wattage ATX 3.1 unit with a native 12V-2x6 connector. Sizing the PSU comfortably above the summed component draw, rather than at the bare minimum, is what prevents transient-trip shutdowns under bursty AI and gaming loads.
Is a 240mm AIO enough, or do I need a 360mm?
A quality 240mm AIO such as the MasterLiquid ML240L handles a mainstream 8-core CPU comfortably, and the GPU manages its own cooling, so the CPU loop rarely needs to grow to 360mm for inference builds. Step up to a larger radiator only if you pair a very high-core, high-TDP CPU with sustained all-core workloads like fine-tuning or heavy compilation alongside the AI tasks.
How much RAM should the build have?
System RAM matters most when you offload model layers off the GPU, so pair the build with a dual-channel kit sized to hold whatever won't fit in VRAM plus your OS and tooling overhead. For pure in-VRAM inference the RAM requirement is modest, but builders who intend to run oversized models with CPU offload should prioritise capacity and matched dual-channel speed over raw core count.
Can I start on an RTX 3060 and upgrade later?
Yes, and it's a sensible on-ramp. A 12GB RTX 3060 runs 7B-to-14B models fully in VRAM today, and because the surrounding CPU, SSD, PSU and cooler carry over, you can drop in a flagship card later without rebuilding. Size the PSU and case for the future GPU up front so the only later change is the graphics card itself.

Sources

— SpecPicks Editorial · Last verified 2026-07-16

Ryzen 7 5800X
Ryzen 7 5800X
$217.45
View price →

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →