Skip to main content
Cooling a 24/7 Local LLM Rig: Air vs 120mm AIO vs 240mm AIO

Cooling a 24/7 Local LLM Rig: Air vs 120mm AIO vs 240mm AIO

Sizing a cooler for continuous inference duty — why the always-on box has different thermal rules than the gaming box.

Local LLM boxes running around the clock face a sustained thermal load, not a burst load. Air coolers usually win — here's when a 120mm or 240mm AIO is actually the right call for a 24/7 rig.

Short answer: For a home-office box running a local LLM around the clock, a good dual-tower or single-tower air cooler like the Noctua NH-U12S is the right default. It has no wet failure mode, no pump MTBF to worry about, and easily dissipates the sustained 100-140W that a Ryzen 7-class CPU pushes during prefill on a modern quantized model. Step up to a 120mm AIO like the NZXT Kraken M22 only in a mini-ITX chassis where a tall heatsink will not physically fit, and step up to a 240mm AIO like the CoolerMaster ML240L RGB V2 only if you are batching heavy inference on a 105W-class CPU in a case with good radiator mount points.

Why gaming cooler reviews mislead here

Almost every published CPU cooler review is a gaming test. That means a five-to-fifteen-minute burst load — the length of a single benchmark run or a boss fight — and a rest period between runs while the reviewer changes titles or resets a scene. Package power spikes to the CPU's boost ceiling for tens of seconds at a time and then falls back as the workload becomes GPU-bound. A cooler that handles those bursts cleanly is scored well; a cooler that lets the CPU throttle after ten seconds is scored badly.

An always-on inference rig is the opposite kind of thermal problem. If you have moved from a cloud API to a local model to cut recurring AI API bills, your box is doing prompt prefill or token generation somewhere between "several times an hour" and "constantly," depending on how heavily you use it. Even a lightly-loaded home-office instance easily accumulates hours per day of near-peak CPU activity, and the way heat behaves under that kind of sustained load is nothing like a five-minute gaming benchmark. Package temperatures reach thermal equilibrium against the case airflow, not against the cooler's peak dissipation. Fan curves get pinned at their steady-state ramp, not their transient one. Pump wear on an AIO accumulates in real hours. Every one of those things flips the ranking a gaming reviewer would produce for the same coolers.

The framing this guide uses is the framing you should be using: this is a sustained-load thermal problem for a machine you are not going to babysit, not a burst-load problem for a machine you are actively driving. The right answer is the cooler that gives you the lowest maintenance risk and the flattest noise profile over three years of 24/7 duty, not the one with the lowest peak temperature in a five-minute benchmark.

Key takeaways

  • Default pick: dual-tower or premium single-tower air (the NH-U12S is the reference here). Zero pump risk, quiet, cheap, and enough for any 105W-class CPU under sustained load.
  • 120mm AIO: only for SFF and mini-ITX where a tall heatsink physically will not fit. Do not choose a 120mm AIO in a case that could accept a tower cooler — the AIO adds a failure mode without giving you meaningful headroom.
  • 240mm AIO: only if you are batching heavy prefill on a 105W-class CPU in a case with a real 240mm radiator mount, and you have accepted the pump-wear tradeoff. Otherwise, an air cooler is fine.
  • Case airflow matters more than the cooler on the CPU once you are past the "big enough" threshold. A good cooler in a stagnant case loses to a mediocre cooler in a well-ventilated one.
  • Noise floor is the design constraint for a rig sharing a room with you. Aim for sustained sub-30 dBA at one meter.

Step 0 — diagnose your actual thermal load

Before you buy a cooler, understand which side of your rig is doing the work. Local LLM inference decomposes into two very different phases, and each puts a very different load on the machine:

Prompt prefill is the phase that ingests the input tokens and builds the initial KV cache. It is heavily parallel and, on any modern quantized runtime, is compute-bound on whatever silicon is doing the math. On a CPU-only or CPU-heavy setup, prefill pushes the package to sustained near-full utilization for however long it takes to chew through the prompt. On a GPU-first setup, prefill is mostly on the GPU, and the CPU stays modest.

Token generation is the phase that produces one token at a time, sampled from the model's logits. It is memory-bandwidth-bound rather than compute-bound. On a GPU-first setup, the CPU is largely idle; on a CPU setup with offload, the CPU stays busy but at lower absolute power than during prefill.

The practical implication: if you are running a 7B or 13B model entirely on a 12GB GPU like the RTX 3060 12GB, your CPU will not be the bottleneck and its cooler load will be modest. If you are running a 30B+ model with CPU offload, or if you are running smaller models entirely on CPU, your CPU is the hot component and its cooler needs to be sized for continuous duty. Match the cooler to the workload, not to the CPU's nameplate TDP.

Spec-delta table

CoolerRadiator / HeatsinkRated fan noiseSocket supportTypical street price
Noctua NH-U12S158mm single-tower, 120mm NF-F1222.4 dBA @ 1500 RPMAM4, AM5, LGA1200/1700$70
NZXT Kraken M22120mm rad + block/pump21-36 dBAAM4, LGA1200/1700$110
CoolerMaster ML240L RGB V2240mm rad + block/pump6-30 dBAAM4, AM5, LGA1200/1700$85
AC Infinity AIRCOM S7Case/cabinet exhaust unit18-32 dBAN/A (external)$80

What sustained load does inference actually put on the CPU?

On a Ryzen 7 5800X (105W nameplate TDP, ~140W PPT ceiling) running a mid-sized model with any meaningful CPU workload, expected package power under sustained load looks roughly like:

  • Idle / KV-cache-only serving: 25-40W. Trivial thermal load.
  • Token generation with partial CPU offload: 60-90W. A stock cooler starts sounding tired but keeps up.
  • Prompt prefill on a long context, CPU-heavy path: 110-140W. This is the p95 you need to size for.
  • Continuous batch prefill (rare on a personal rig, common on a homelab serving multiple users): 130-140W sustained for minutes at a time.

The reason p95 matters more than peak: a five-second peak is trivially absorbed by any cooler with mass. A five-minute sustained 130W load reaches thermal equilibrium against the cooler and the case airflow, and that equilibrium temperature is what determines whether your CPU throttles, how loud the fan gets, and how much heat is dumped into the rest of the case. Size for the p95, not for the peak marketing number.

Air cooling the always-on box — the Noctua NH-U12S case

The strongest argument for air cooling a 24/7 rig is not thermal performance — a 240mm AIO will beat a single-tower air cooler under sustained peak load, and there is no honest way around that. The strongest argument is failure mode. A Noctua NH-U12S is a lump of aluminum with a fan on it. There is no pump, no sealed liquid loop, no PWM header for pump speed, and no possibility of a leak. When the fan bearing eventually reaches end of life — typically 6+ years of continuous use per Noctua's own MTBF ratings — it gets louder before it fails, giving you months of warning. If the worst case does happen and the fan seizes, the CPU throttles down instead of catching fire.

An AIO fails differently. Pump wear accumulates in real hours; a pump rated for 60,000 hours has burned roughly 26,000 of those in three years of 24/7 use. When it fails, it usually fails quietly — the pump stops circulating but the fans keep spinning, and unless you are watching the package temperature graph you may not notice until the CPU thermal-throttles under load. In the worst case, a failing pump loses coolant to the case, which is an inconvenience on a home rig and a disaster on a rack rig.

The NH-U12S is 158mm tall, which fits in every mid-tower and most mini-tower cases; verify against your motherboard for RAM clearance if you are running tall RGB memory kits, but the U12S was designed with clearance in mind and generally clears standard-profile DIMMs without issue. It comes with mounting hardware for AM4 (native for the 5800X), AM5, LGA1200, and LGA1700. For most 105W-class CPUs, it delivers sub-80°C package temps under sustained 130W load with the included NF-F12 fan running at 1200-1400 RPM, comfortably below 25 dBA at one meter.

When does a 120mm AIO make sense?

Only in physical-clearance cases. The NZXT Kraken M22 exists to solve a specific problem: small-form-factor and mini-ITX chassis where a 158mm-tall tower cooler simply will not fit under the side panel. In an SFF NCASE M1 or an NR200P, a Node 202, or a shoebox-style ITX build meant to sit next to a router, a 120mm AIO gives you liquid-cooler-class dissipation in a form factor that a tower cannot occupy.

Outside of that clearance problem, a 120mm AIO is a bad choice for a 24/7 rig. It adds a pump failure mode, adds pump noise (M22 pumps have a specific whine some people notice), and delivers thermal performance roughly comparable to a good single-tower air cooler — you are trading down on reliability for no thermal win. If your case can accept the NH-U12S, choose the NH-U12S.

When do you actually need a 240mm?

If you are batching heavy prefill on a 105W-class CPU in a case with a real 240mm radiator mount, and you have accepted the pump-wear tradeoff, a 240mm AIO like the CoolerMaster MasterLiquid ML240L RGB V2 gives you meaningful headroom over air. Under sustained 140W, a 240mm AIO holds a Ryzen 7 5800X 5-10°C cooler than the U12S, which translates to a slightly quieter fan curve at the same package temperature or a slightly higher boost residency at the same fan speed.

The ML240L V2 is the pragmatic pick in this tier — it is inexpensive, uses a competent Third Generation pump, and pairs with reasonable stock fans. Higher-end 240mm options exist (Arctic Liquid Freezer II 240, EK Nucleus, Corsair H100i) but the price step gives you diminishing returns on a rig where the cooler is one of six components and you are not chasing enthusiast-grade overclocks.

Do not pair a 240mm AIO with a cheap case that has no dedicated 240mm mount point. Bolting a radiator onto a case that is not designed to accept it produces bad airflow, awkward tube routing, and often the radiator sitting in the intake or exhaust of the front fans in ways that hurt overall system cooling.

Case airflow is the other half

The single most-overlooked variable in a 24/7 build is case airflow. A rig running a modern GPU like the ZOTAC RTX 3060 Twin Edge is dumping around 170W of heat into the chassis under continuous inference load. If that heat cannot leave the case, it raises ambient temperature inside the box, and every component — CPU, GPU, VRMs, NVMe SSD, memory — runs hotter permanently. A cooler is only as good as the ambient air it has access to.

The design points to hit:

  • Positive pressure, with intake fans running slightly faster than exhaust, to keep dust out of the case.
  • A clear front-to-back path for the GPU exhaust to leave the case without recirculating into the CPU cooler intake.
  • At least two 120mm or 140mm intake fans at the front, one 120mm exhaust at the rear, one 120mm or 140mm exhaust at the top for any rig with a card drawing over 150W.

If your rig lives in a closed cabinet or on a rack shelf, a case-airflow strategy alone is not enough — the cabinet itself becomes the thermal enclosure. The AC Infinity AIRCOM S7 is the standard answer here: a top-mounted 12-inch exhaust unit that pulls warm air out of a media cabinet or a networking closet. Pair it with a temperature-triggered fan controller and you can hold a cabinet ambient within 5°C of the room, which is the difference between a rig that runs quietly for years and a rig that thermal-throttles every summer afternoon. For a full airflow breakdown, see our Best CPU Cooler for the Ryzen 7 5800X: Air vs AIO in 2026 walkthrough.

Noise floor at 3am

The number that matters is sustained dBA at one meter. A rig that hits 40 dBA for two minutes during a benchmark run is unremarkable. The same 40 dBA running through the night from a rig sharing your home office is intrusive within a week.

The design targets for a home-office 24/7 build:

  • Sustained under 30 dBA at one meter for a machine sharing a room with you.
  • Sustained under 25 dBA if it shares a bedroom or an open-plan space where audible noise is a real disruptor.
  • Fan curves tuned for a flat response, not a reactive ramp — you want the fans to hold a modest RPM continuously, not to spin up and down in ten-second cycles that draw attention.

Air coolers with 120mm or 140mm fans hit these numbers more easily than radiator-based AIOs at equivalent thermal load, because the larger fan area lets you move the same volume of air at a lower RPM. This is the second big reason air wins for always-on boxes: it is quieter at the same steady-state temperature, not just cheaper and more reliable.

Benchmark table

Representative values under a sustained CPU-heavy inference workload on a Ryzen 7 5800X in a well-ventilated mid-tower case at 22°C ambient, sourced from published cooler testing at Tom's Hardware and Puget Systems:

CoolerPackage tempFan RPM (steady)Measured dBA @ 1m
Noctua NH-U12S74°C125024
NZXT Kraken M22 (120mm AIO)76°C140032
CoolerMaster ML240L V2 (240mm AIO)68°C110027
Stock Wraith Prism (reference)92°C (throttling)260043

The 240mm AIO wins on absolute temperature, but the noise floor tells a different story — the U12S delivers within 6°C of the ML240L at almost 3 dBA lower sustained noise, and without a pump. For a 24/7 box, that is the trade you want to make.

Perf-per-dollar and failure-risk math

Air cooler: $70 upfront, essentially zero risk of catastrophic failure, fan replacement is a $20 job when it eventually happens, expected lifetime well over five years of 24/7 use.

120mm AIO: $110 upfront, pump MTBF typically rated in the 50,000-70,000 hour range, three years of 24/7 use consumes roughly 26,000 hours of that rating, replacement is a full unit swap.

240mm AIO: $85 upfront for the ML240L V2, same pump-wear math as the 120mm AIO, better thermal headroom for the same wear budget, more physical bulk to fit in the case.

Amortized over three years, the cost delta is small in absolute terms. The reason we default to air for a 24/7 rig is not cost — it is that "essentially zero risk" of the air cooler beats "small but non-zero risk" of the AIO on a machine you are not going to actively babysit.

Verdict matrix

Get air if: Your case can accept a 158mm-tall tower cooler, your CPU is a 65W or 105W-class part, and you value zero-pump-risk over the 5-10°C of thermal headroom a 240mm AIO gives you. This is the default for a home-office 24/7 rig and covers the majority of readers.

Get a 120mm AIO if: You are building in a mini-ITX chassis where a tower cooler physically will not fit, and you have accepted the pump-wear tradeoff in exchange for the space savings. Do not choose a 120mm AIO in a case that could accept a tower cooler.

Get a 240mm AIO if: You are batching heavy inference workloads on a 105W-class CPU, your case has proper 240mm radiator mount points, you want the last 5-10°C of package temperature headroom, and you have accepted a three-year pump-wear commitment as an acceptable maintenance line.

Bottom line

For a 24/7 local LLM box on a Ryzen 7 5800X or an equivalent 105W-class part in a standard mid-tower, buy the Noctua NH-U12S and spend the difference on case fans. It handles the sustained inference load with margin, it will not fail in a way that damages the rest of your build, and it is quieter than any AIO you can buy at three times the price. If your rig lives in a closed cabinet or on a rack shelf, add an AC Infinity AIRCOM S7 to move the cabinet's ambient air, because no CPU cooler solves an enclosure problem. Only reach for a 240mm AIO like the CoolerMaster ML240L V2 if you are doing heavy batch prefill and you have accepted the pump-wear tradeoff for the extra thermal headroom.

Related guides

Citations and sources

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Tap any product for full specs, live Amazon & eBay pricing, and alternatives.

SpecPicks earns a commission on qualifying purchases through both Amazon and eBay affiliate links. Prices and stock update independently.

Watch a review

Friendly Fire: AMD Ryzen 7 5800X CPU Review & Benchmarks vs. 5600X & 5900X — Gamers Nexus on YouTube

Frequently asked questions

Does a local LLM even load the CPU enough to need a good cooler?
It depends on which phase dominates your usage. Token generation on a GPU leaves the CPU largely idle, but prompt prefill, embedding runs, CPU-offloaded layers and any RAM-resident portion of a large model push the package toward sustained near-full utilization. On a 105W-class part like the Ryzen 7 5800X, that is a continuous load rather than the seconds-long burst a gaming benchmark measures, and stock coolers throttle under it.
Are AIO pumps reliable enough to run continuously for years?
Modern sealed AIO pumps are rated for tens of thousands of hours, but a pump is an additional wear component with a wet failure mode that a heatsink simply does not have. On a machine running every hour of every day, a three-year duty cycle consumes a meaningful share of that rating. Air coolers fail gracefully — a bearing gets loud long before it stops — which is why air remains the default for always-on builds.
How loud is too loud for a rig in a home office?
Sustained noise is judged very differently from intermittent noise. A gaming rig hitting 40 dBA during a session is unremarkable, while the same 40 dBA running through the night becomes intrusive. Target a sustained sub-30 dBA at one meter for a machine sharing a room with you, which in practice means large slow fans, a generous heatsink or radiator, and a fan curve tuned for a flat response rather than a reactive ramp.
Should I cool the GPU differently for inference than for gaming?
The GPU sees the inverse pattern of the CPU: near-continuous load during generation rather than bursts. Practically, that means case airflow matters more than the cooler on the card itself. Ensure the card has unobstructed intake, keep at least one slot of clearance below a triple-fan design, and prioritize front-to-back exhaust. Undervolting typically sheds meaningful heat and noise at a small throughput cost for continuous workloads.
Can I put an always-on inference box in a closed cabinet?
Only with forced ventilation. A closed cabinet with a continuously loaded machine inside reaches thermal equilibrium well above ambient, and every component then runs at that elevated baseline permanently. An external cabinet fan system that exhausts warm air and draws cool air in is the minimum requirement. Without one, expect sustained clocks to fall and component lifetime to shorten, regardless of which cooler is on the CPU.

Sources

— SpecPicks Editorial · Last verified 2026-08-09

Ryzen 7 5800X
Ryzen 7 5800X
$219.00
View price →

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →