Skip to main content
Best GPU for a NAS or Homelab Running Local LLMs (2026)

Best GPU for a NAS or Homelab Running Local LLMs (2026)

The 12GB tier is the practical floor for an always-on box — and chassis depth, PSU connectors and idle watts decide the pick long before compute does.

A 12GB card is the practical floor for homelab inference. Sourced tok/s, VRAM tiers, and the chassis and PSU checks to run before you buy anything.

Quick answer: which GPU for which model, at Q4

Start from the model you want to run. At Q4_K_M the weights take roughly 0.55 GB per billion parameters and the runtime plus a usable context window wants about 2 GB on top, which is where the VRAM column comes from. The card in each row is the cheapest one SpecPicks holds measurements for that clears that figure — not the fastest — and every tokens-per-second number is a median over community-reported Q4 runs, with the run count and the source beside it.

Model you want to run VRAM you need Cheapest card that clears it Measured Price Source
7-9B (Llama 3.1 8B, Qwen 3 8B)~5 GB of weights at Q4 8 GBweights plus a usable context window NVIDIA GeForce RTX 30508 GB 28.9 tok/s8 runs · 3 sources $249MSRP llmrun.dev
12-14B (Qwen 3 14B, Phi-4)~8 GB of weights at Q4 12 GBweights plus a usable context window Arc B58012 GB 35 tok/s4 runs · 4 sources $310street llama.cpp GitHub Discussions
20-27B (Gemma 3 27B, Mistral Small)~15 GB of weights at Q4 16 GBweights plus a usable context window NVIDIA GeForce RTX 4070 Ti SUPER16 GB 88.7 tok/s8 runs · 2 sources $830street LocalLLaMA
30-35B (Qwen 3 32B, QwQ 32B)~19 GB of weights at Q4 24 GBweights plus a usable context window NVIDIA GeForce RTX 309024 GB 29.2 tok/s6 runs · 6 sources $1,550street GitHub (thc1006)
70B+ (Llama 3.3 70B, Qwen 2.5 72B)~40 GB of weights at Q4 48 GBweights plus a usable context window AMD Radeon Pro W7900 48GB48 GB 11.3 tok/s14 runs · 3 sources $3,999MSRP Windows Forum (citing AMD)

Best AI rigs: the same table with the fastest card in each band → Full local-LLM GPU buying guide How we source these numbers

Hardware at a Glance

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Street price is the lowest tracked listing within a sane band of MSRP; prices move daily. Rows marked for comparison are not covered by this article — they are the nearest cards by VRAM, included so the throughput column has something to be read against.

GPUVRAM Llama-3-8B class, Q4Price Source
NVIDIA GeForce RTX 3060 12 GB 55.2 tok/s25 runs · 13 sources $329MSRP smeltcore.com
NVIDIA GeForce RTX 5070for comparison 12 GB 59.1 tok/s5 runs · 5 sources $501street knightli.com
Arc B580for comparison 12 GB 40 tok/s19 runs · 12 sources $310street llama.cpp GitHub Discussions

As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices shown are indicative and may vary. See our editorial methodology.

For a NAS or homelab box that also has to run local LLMs, a 12GB card is the entry point that actually clears the bar: an RTX 3060 12GB generates roughly 64.5 tok/s on Llama 3.1 8B at q4_K_M with an 8K context, per TyoLab's 12GB VRAM benchmark, on a card TechPowerUp rates at 170 W board power. VRAM, not compute, is the ceiling that decides the pick.

Who this guide is for

You already have a box that never turns off. It runs Jellyfin or Plex, maybe Immich, maybe a Proxmox host with a handful of LXC containers and a TrueNAS VM holding the family photo archive. It has a case that was chosen for drive bays rather than GPU clearance, a power supply sized for spinning disks, and a spot in a closet or a rack where noise matters more than frame rates. And now you want to add local inference to it — an Ollama endpoint your laptop can hit, a summarizer for the RSS pile, an embedding job for a document search index, maybe a coding assistant that never phones home.

This is a genuinely different problem from "which GPU should I buy for AI." The buying guides that answer that question optimize for peak tokens per second and assume a 750 W tower with three free slots. A homelab GPU decision is dominated by four constraints that those guides ignore: the physical envelope of a chassis designed around drive cages, the PCIe power connectors your existing PSU actually has, the wattage penalty of a card that idles 8,760 hours a year, and the VRAM ceiling that decides which model classes are even reachable.

The good news is that those constraints converge on a fairly narrow answer, and it is cheaper than the AI-rig discourse suggests. The 12GB tier is the sweet spot in 2026 because it is the smallest pool that holds a useful 8B-to-14B model at a quantization you can trust, with enough headroom left over for a KV cache at real context lengths. Below that you are stuck with 3B-class models. Above it you pay a steep premium for capability most homelab workloads do not use.

What follows works from the constraint backwards — chassis, power, VRAM, model class — rather than from a leaderboard forwards. Every figure is sourced. Where a number depends too heavily on a specific build to state honestly, that is said plainly instead of guessed.

Key takeaways

  • 12GB is the practical floor for a homelab inference box in 2026. It hosts 8B-class models fully resident with room for a long context, and 12B–14B models at q4 with measured VRAM use around 8.1–9.5 GB per llmrun.dev's RTX 3060 12GB data.
  • Idle watts matter more than peak watts. A homelab card spends most of its life waiting for a prompt. Peak board power is a rated ceiling (170 W on the RTX 3060), not an operating cost.
  • Measure the chassis before you buy. Triple-fan 12GB cards run past 300 mm long and eat more than two slots. Many NAS and cube cases will not take one regardless of how good the card is.
  • Your PSU needs a real PCIe 8-pin, not just a spare EPS lead. NVIDIA's published system requirement for the RTX 3060 class is a 550 W supply (NVIDIA product page).
  • Model class follows VRAM, not brand. 8GB gets you 3B–8B at tight context. 12GB gets you 8B comfortably and 12B–14B at q4. 16GB+ opens 22B-class models at low quantization.
  • Storage is a cold-start problem, not a throughput problem. Weights load once; a SATA SSD turns a minute-long swap into a few seconds.

Step 0 — diagnose the constraint before you buy anything

Work this decision tree in order. The first "no" is your real ceiling, and it is usually not the one people shop for.

1. Does a dual-slot, 300 mm card physically fit? Open the case and measure from the PCIe bracket to the nearest drive cage, and from the motherboard surface to the side panel. If the answer is under about 270 mm of length or under two full slots of height, no amount of budget fixes it — you are looking at a low-profile or single-fan card, or an eGPU enclosure, or a CPU-only build.

2. Does the PSU have a PCIe 8-pin? Many NAS-oriented and SFF supplies ship with CPU EPS connectors and SATA power only. They are not interchangeable, and adapters off Molex chains are exactly the kind of thing that fails at 3 a.m. under load. Check for the connector before you check the wattage.

3. Can the box tolerate the 24/7 wattage delta? Estimate the annual cost as idle watts × 8,760 hours × your kWh rate, then add the load hours you will actually generate. This number surprises people in both directions: it is smaller than the peak-power headline suggests, and larger than "it's just a GPU" assumes.

4. Only now — how much VRAM do you need? This is the fun part, and it is the last question, not the first.

If steps 1 through 3 pass, the rest of this guide is a spec conversation. If any of them fail, skip to the "when to stay CPU-only" section at the end, because the honest answer for your build may be no GPU at all.

How much VRAM does a homelab LLM box actually need?

VRAM sets the hard boundary on which models are reachable at usable speed. Everything else — CUDA core count, memory bandwidth, clock — decides how fast you traverse a model that already fits. Cross the ceiling and layers spill to system RAM, and throughput does not degrade gracefully; it falls off a cliff.

VRAM tierModel classes that fit fully residentTypical q4 context headroomWhat it feels like
8 GB1B–8B at q4Short — 4K comfortable, 8K tightFine for summarization and routing; 8B models fit but leave little KV cache room
12 GB1B–8B comfortably, 12B–14B at q48K–16K on 8B, 4K–8K on 14BThe homelab sweet spot; interactive chat plus a second workload
16 GBUp to ~22B at low quantizationGenerous on 8B, workable on 14BOpens the 22B tier, but at q3-class quality tradeoffs
24 GB24B–32B at q4Long contexts on everything below 24BRemoves the compromise; costs roughly double

The measured VRAM figures behind the 12GB row come from llmrun.dev's RTX 3060 12GB page: Gemma 4 12B at q4_K_M uses 8.2 GB and generates 28.4 tok/s, Mistral NeMo 12B at q4_K_M uses 8.1 GB at 29.0 tok/s, Llama 2 13B at q4_K_M uses 8.6 GB at 27.2 tok/s, and Phi-4 14B at q4_K_M uses 9.5 GB at 24.6 tok/s. That is a coherent picture: a 12B–14B model at q4 lands in the 8–9.5 GB band and leaves 2.5–4 GB for the KV cache and framebuffer overhead.

Go one tier up in parameters and the picture breaks. A 26B-class model at q4_0 on a 12GB-class card is reported at 5.0 tok/s in community measurements collected from r/LocalLLaMA — roughly a sixth of the speed of the 14B tier, because the model no longer fits and the offloaded layers are executing on the CPU across the PCIe bus.

Which GPU fits a NAS chassis and a 24/7 power budget?

The reference pick for this build is the MSI GeForce RTX 3060 Ventus 3X 12G OC. It is a 12GB GDDR6 Ampere card that TechPowerUp lists at 170 W board power with 3,584 CUDA cores, and NVIDIA specifies a 550 W system power supply for the class. It takes a single PCIe 8-pin. That last detail is the one that decides whether your existing NAS PSU is adequate or whether the project just grew a power-supply purchase.

The catch is physical. Ventus 3X is a triple-fan design; MSI's own dimensions put it past 300 mm long and over 50 mm thick, which is a 2.5-slot footprint in practice even where the spec sheet says two. In a mid-tower with the drive cage removed, that is a non-event. In a four-bay cube NAS or a 2U rackmount, it is a wall. Measure first, and if you are short on length, look at dual-fan and blower variants of the same GPU rather than dropping to the 8GB tier — the silicon is what matters, and the 12GB pool is the whole reason you are here.

Why not the 8GB tier, which is cheaper and shorter? Because the model classes it hosts are the ones you were already running acceptably on CPU. The 8GB tier does 3B and 8B at short context. The moment you want a 12B–14B model, or 16K of context on an 8B model, you are back to offload. The 12GB card is the cheapest way to stop thinking about the ceiling.

Why not go straight to 16GB or 24GB? You can, and if the budget is there it removes future regret. But the homelab workload profile — a few interactive chats a day, some batch summarization overnight, an embedding job — does not saturate the extra capacity, and the extra card costs roughly what the rest of the NAS upgrade did.

Quantization matrix: what actually fits in 12GB

Quantization is the lever that decides whether a model is resident. These are approximate GGUF weight sizes for an 8B-class model as published in the Ollama model library, paired with measured throughput on 12GB Ampere hardware where a sourced figure exists.

QuantApprox. weights (8B-class)Fits in 12GB with context?Measured tok/s referenceQuality notes
q2_K~3.2 GBYes, with lots of roomNot separately measuredNoticeable degradation; use only when nothing else fits
q3_K_M~4.0 GBYes13.6 tok/s on a 22B model at q3_K_S, 11.2 GB used (Local AI Master)Acceptable for routing, weak for reasoning
q4_K_M~4.9 GBYes, the default choice64.5 tok/s Llama 3.1 8B @ 8K (TyoLab)The quality/size knee; what most people should run
q5_K_M~5.7 GBYesNot separately measuredMarginal quality gain over q4_K_M for ~16% more VRAM
q6_K~6.6 GBYes, tighter contextNot separately measuredNear-lossless for most tasks
q8_0~8.5 GBYes, but context gets tight211.9 tok/s on a 1B model @ 8K (TyoLab)Effectively lossless; rarely worth the VRAM on 12GB
fp16~16 GBNoExceeds the pool before any KV cache

The practical read: q4_K_M is the setting a 12GB homelab card should default to. It leaves the most room for context, and the measured throughput at that setting — 64.5 tok/s on Llama 3.1 8B, 59.5 tok/s on Qwen 3 8B, 66.2 tok/s on a DeepSeek-R1 7B distill, all per TyoLab — is comfortably faster than reading speed.

Prefill vs generation: why RAG stresses the half nobody benchmarks

Almost every tokens-per-second number you see quoted is generation throughput — how fast the model emits new tokens once it has read your prompt. Prefill is the other half: how fast it ingests the prompt in the first place. They are different operations with different bottlenecks, and homelab workloads lean on prefill far harder than chat benchmarks suggest.

The numbers are not close. LocalScore's RTX 3060 accelerator page records 1,490 tok/s prefill against 52.2 tok/s generation on Llama 3.1 8B at q4_K_M, and 6,064 tok/s prefill against 184 tok/s generation on a 1B model. llama.cpp discussion #10879 records 1,815.7 tok/s prefill and 75.94 tok/s generation for a 7B model at q4_0 on the same class of card. Prefill runs roughly 20–30× faster than generation because it is a compute-bound batched operation rather than a memory-bandwidth-bound sequential one.

That ratio is why RAG changes the calculus. A retrieval pipeline that stuffs 8,000 tokens of retrieved documents into every prompt is asking the card to prefill 8,000 tokens before it emits the first word of the answer. At 1,490 tok/s that is about 5.4 seconds of dead air. Reduce the retrieved-chunk budget and that latency drops linearly. Add a second concurrent request and it stacks.

The homelab implication: if your primary use is document Q&A over a personal archive, prefill throughput and context budget matter more than the headline generation figure, and cutting retrieved context is a bigger latency win than a faster card.

Context-length impact: where 12GB actually runs out

The KV cache grows linearly with context length, and it comes out of the same VRAM pool as the weights. This is the mechanism by which a model that "fits" stops fitting.

Hardware Corner's RTX 3060 12GB measurements show the progression cleanly on Qwen 3 8B at q4_K_XL:

ContextVRAM usedPrefill tok/sGeneration tok/s
4K6.0 GB1,696.855.2
16K7.5 GB1,119.242.0
32KNot separately reported31.9

Generation throughput falls about 42% going from 4K to 32K context, and VRAM use climbs 1.5 GB between 4K and 16K on an 8B model. Extrapolate that to a 14B model already using 9.5 GB at short context and the ceiling arrives fast: a 14B model at 16K context is at or past the edge of a 12GB pool.

The 32K figure of 31.9 tok/s (via singhajit.com's compilation) is still interactive. That is the useful conclusion — a 12GB card at 32K context on an 8B model is slower but not broken. The same context on a 14B model is where you start offloading.

Does the host CPU matter?

If the whole model is resident in VRAM, barely. The CPU handles tokenization, sampling glue, and the HTTP server, and any modern eight-core part is more than adequate. The moment you exceed VRAM, the CPU and its memory bandwidth become the throughput ceiling, because the offloaded layers execute there.

The AMD Ryzen 7 5800X is the sane 24/7 host for this build. PassMark records it at a 27,661 CPU Mark with a 3,448 single-thread rating, 8 cores and 16 threads at 3.8 GHz base and 4.7 GHz boost, on a 105 W TDP and the AM4 socket (PassMark). Sixteen threads is generous headroom for a box that is also running containers, and AM4 boards and DDR4 are cheap in 2026 in a way that AM5 is not.

The cheaper path is a second-hand LGA1151 platform around the Intel Core i7-9700K, which PassMark rates at a 14,400 CPU Mark and 2,857 single-thread, 8 cores and 8 threads at 3.6 GHz base and 4.9 GHz turbo on a 95 W TDP (PassMark). It gives up roughly 48% of the multithreaded score and 17% of the single-thread score to the 5800X, which matters when layers are offloaded and matters very little when they are not. If your model always fits in VRAM and the board is already paid for, the 9700K host is not the bottleneck.

Dual-channel memory bandwidth is the variable that actually moves offload throughput, more than core count. If you expect to run models that spill, prioritize populating both channels with the fastest RAM the platform supports over adding cores.

Where do the model weights live?

Weights are read once into VRAM or system RAM and then never touched again until you swap models. Storage therefore does not affect generation throughput at all — it affects cold-start latency, and only that.

It affects it a lot, though. A q4_K_M 8B model is roughly 5 GB on disk; a library with six quantizations of three models is comfortably 60–100 GB. Loading 5 GB from a spinning disk at 120 MB/s is about 42 seconds. Loading it from a SATA SSD at 500 MB/s is about 10 seconds. If you rotate models — and everyone who runs a homelab inference box rotates models — that difference is the whole user experience.

Two cheap SATA tiers cover this:

Neither is a fast drive by NVMe standards, and that is fine. Both saturate what a cold model load needs. Put the models on solid state and the OS wherever you like; do not put them on the array.

Keeping it quiet and cool at 24/7 duty

Two heat sources, two different problems.

The CPU runs continuously at low load with occasional bursts, which is the profile air cooling handles best. The Noctua NH-U12S is the standard answer for a quiet 105 W-class host: a 120 mm single-tower design that fits under most side panels and clears tall RAM. On a 5800X in a homelab duty cycle it never gets loud, because it is almost never being asked to dissipate 105 W.

The GPU is the harder one, and the problem is airflow rather than cooler capacity. A triple-fan open-air card is designed to dump heat into a case with front-to-back flow. NAS chassis are designed to pull air across drive cages, which is not the same path. The failure mode is not the GPU overheating — Ampere cards throttle gracefully — it is the GPU exhaust raising the ambient temperature around the drives.

Two mitigations worth the effort: leave a slot of clearance below the card so the bottom fan is not breathing against a PSU shroud, and set a conservative fan curve rather than the aggressive default, since a homelab card at 30% load does not need 60% fan speed. If the case genuinely has no room for air to move, a blower-style card is the correct tradeoff — louder under load, but it exhausts out the back instead of into the drive bays.

Multi-GPU in a homelab: when a second 12GB card beats one 24GB card

The appeal is obvious: two 12GB cards are 24GB of VRAM for less than one 24GB card. The reality has conditions.

Two 12GB cards win when you are running two independent workloads — say, an inference endpoint on one card and Jellyfin NVENC transcoding on the other. Isolation is genuinely valuable here, and it is the case the homelab actually presents. It also wins when you want to serve two model sizes concurrently without swapping.

One 24GB card wins when you want to run a single model larger than 12GB. Splitting a model across two cards over PCIe works in llama.cpp and vLLM, but the interconnect between two consumer cards is the PCIe bus, not NVLink, and the per-layer transfer cost eats much of the theoretical gain. A 24B model split across two 12GB cards is slower than the same model on one 24GB card by a wide margin.

Neither wins when your chassis has one usable slot, your PSU has one PCIe 8-pin, or your case has no airflow path for two open-air cards stacked adjacent. Which describes most NAS builds. Check the physical constraints before you plan around a second card.

Perf-per-dollar and perf-per-watt

Two ratios that matter for an always-on box, computed from the sourced figures above. Street prices are indicative as of 2026 and vary; the 3060 12GB launched at a $329 MSRP per TechPowerUp.

ConfigurationModel / quantMeasured gen tok/sBoard powertok/s per 100 W
RTX 3060 12GBLlama 3.1 8B q4_K_M @ 8K64.5170 W rated~37.9
RTX 3060 12GBQwen 3 8B q4_K_XL @ 16K42.0170 W rated~24.7
RTX 3060 12GBPhi-4 14B q4_K_M24.6170 W rated~14.5
RTX 3060 12GB26B-class q4_0 (offloading)5.0170 W rated~2.9

Throughput figures per TyoLab, Hardware Corner, llmrun.dev and r/LocalLLaMA; board power per TechPowerUp.

The shape of that table is the whole argument. Efficiency collapses by an order of magnitude the moment the model stops fitting. Staying inside the VRAM budget is worth more than any amount of extra compute.

Note that these use rated board power, not measured draw. A homelab card generating for ten minutes a day and idling the rest draws nowhere near 170 W on average — the honest annual figure is dominated by idle, which varies by board partner, driver version, and how many displays are attached. Measure yours at the wall rather than trusting a spec sheet.

Common pitfalls

  • Buying on CUDA cores instead of VRAM. A faster 8GB card is worse than a slower 12GB card for this job, every time, because the ceiling is capacity.
  • Assuming a spare PSU cable is a PCIe cable. CPU EPS 8-pin and PCIe 8-pin are keyed differently and carry different rails. Adapters from SATA or Molex chains are a fire-safety compromise, not a solution.
  • Sizing context by what the model supports rather than what fits. A model advertising 128K context on a 12GB card will happily accept the setting and then thrash. Set context to what the KV cache budget allows.
  • Putting the model library on the spinning array. It works, and every model swap costs you 40 seconds you did not need to spend.
  • Ignoring the transcode collision. If the same card serves Jellyfin, NVENC and CUDA coexist on the die but compete for the VRAM pool. A model using 9 GB of 12 leaves thin room for concurrent 4K transcode sessions.

Verdict matrix

Get the 12GB tier if you want an always-on inference endpoint for 8B–14B models, your chassis takes a dual-slot card, and your PSU has a PCIe 8-pin. This is the majority case and the MSI RTX 3060 Ventus 3X 12G is the reference implementation.

Get 16GB or more if you specifically want 22B-class models, you already know you will run long contexts on 14B models, or you plan to serve multiple concurrent users. The extra pool removes the compromise rather than deferring it.

Skip the GPU and stay CPU-only if your workloads are batch — overnight summarization, embedding jobs, scheduled classification — where latency does not matter. A modern eight-core host running 3B-to-8B models at a few tokens per second is genuinely adequate for that profile, and you save both the card and the idle wattage. It is also the right answer when the chassis or PSU audit in Step 0 failed: an inference container that is slow is better than a GPU that does not fit.

Bottom line

The configuration that answers this question for most homelabs in 2026: a 12GB RTX 3060 on an AMD Ryzen 7 5800X host with 32GB of dual-channel DDR4, models on a Crucial BX500 1TB SATA SSD, and a Noctua NH-U12S keeping the CPU quiet. Run 8B models at q4_K_M with a 8K–16K context and you get roughly 42–64 tok/s on measured public benchmarks, on a box that is already doing five other jobs.

Buy the VRAM first. Everything else on that list is substitutable.

Related guides

Frequently asked questions

Will a 12GB GPU fit in a typical NAS or small-form-factor server chassis? Not always. Triple-fan 12GB cards run past 300 mm long and occupy roughly 2.5 to 3 slots, which is more than many rackmount and cube NAS chassis allow. Measure clearance from the PCIe bracket to the drive cage before ordering, check that the case takes a dual-slot-plus card, and confirm the power supply has the required 8-pin PCIe connector rather than only a CPU EPS lead.

How much does adding an inference GPU raise my 24/7 power bill? Idle draw matters far more than peak for an always-on box, because a homelab GPU spends most of its life waiting for a prompt. Ampere-class cards idle well below their rated board power and only approach it during generation. Estimate cost as idle watts times 24 hours times your local kWh rate, then add the short bursts of full-load draw for however many minutes per day you actually run inference.

Can I run Jellyfin transcoding and an LLM on the same GPU? Yes, with caveats. NVENC transcoding and CUDA inference use different blocks on the die, so they coexist, but they compete for the same VRAM pool. A 12GB card hosting a quantized model in roughly 8 to 9 GB leaves thin headroom for multiple simultaneous 4K transcodes. Cap concurrent transcode sessions, or pin the model to a smaller quantization so the video pipeline never gets starved mid-stream.

Do I need a powerful host CPU if the GPU does the inference? Only if you plan to offload layers. When the whole model fits in VRAM the CPU mostly handles tokenization, sampling glue, and the API server, and a modest chip is fine. The moment you exceed VRAM and spill layers to system RAM, memory bandwidth and core count dominate throughput, and an eight-core part with fast dual-channel DDR4 becomes the difference between usable and unusable generation speed.

When should I skip the GPU and keep the homelab CPU-only? If your workloads are batch summarization, overnight embedding jobs, or anything where latency does not matter, CPU-only inference on a modern eight-core host is often good enough and saves both the card cost and the idle wattage. Add the GPU when you need interactive chat latency, when you run a coding assistant against the box during the workday, or when prefill on long RAG contexts becomes the bottleneck.

Citations and sources

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Live Amazon & eBay pricing, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

THE INTEL ARC B580 IS ACTUALLY GREAT & AFFORDABLE — Linus Tech Tips on YouTube

Frequently asked questions

Will a 12GB GPU fit in a typical NAS or small-form-factor server chassis?
Not always. Triple-fan 12GB cards like the Ventus 3X run past 300 mm long and occupy roughly 2.5 to 3 slots, which is more than many rackmount and cube NAS chassis allow. Measure clearance from the PCIe bracket to the drive cage before ordering, check that the case takes a dual-slot-plus card, and confirm the power supply has the required 8-pin PCIe connector rather than only a CPU EPS lead.
How much does adding an inference GPU raise my 24/7 power bill?
Idle draw matters far more than peak for an always-on box, because a homelab GPU spends most of its life waiting for a prompt. Ampere-class cards idle in the low tens of watts and only approach their rated board power during generation. Estimate cost as idle watts times 24 hours times your local kWh rate, then add the short bursts of full-load draw for however many minutes per day you actually run inference.
Can I run Jellyfin transcoding and an LLM on the same GPU?
Yes, with caveats. NVENC transcoding and CUDA inference use different blocks on the die, so they coexist, but they compete for the same VRAM pool. A 12GB card hosting a quantized model in roughly 8 to 9 GB leaves thin headroom for multiple simultaneous 4K transcodes. Cap concurrent transcode sessions, or pin the model to a smaller quantization so the video pipeline never gets starved mid-stream.
Do I need a powerful host CPU if the GPU does the inference?
Only if you plan to offload layers. When the whole model fits in VRAM the CPU mostly handles tokenization, sampling glue, and the API server, and a modest chip is fine. The moment you exceed VRAM and spill layers to system RAM, memory bandwidth and core count dominate throughput, and an eight-core part with fast dual-channel DDR4 becomes the difference between usable and unusable generation speed.
When should I skip the GPU and keep the homelab CPU-only?
If your workloads are batch summarization, overnight embedding jobs, or anything where latency does not matter, CPU-only inference on a modern eight-core host is often good enough and saves both the card cost and the idle wattage. Add the GPU when you need interactive chat latency, when you run a coding assistant against the box during the workday, or when prefill on long RAG contexts becomes the bottleneck.

Sources

— Mike Perry · Last verified 2026-09-05

Parts this article names

Amazon Associate — prices tracked 2026-09-05, may vary.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More buying guides from SpecPicks

Browse all buying guides →