As an Amazon Associate, SpecPicks earns from qualifying purchases. See the SpecPicks review methodology.
Buy VRAM first, and buy it now. A 12 GB RTX 3060 runs Qwen2.5 14B Instruct Q4_K_M at 26.6 tokens/s with a 1.92-second time to first token, against roughly 4 tok/s and a 55-second wait on any AM4 CPU (LocalScore). The reason to move this quarter rather than next: TrendForce reported DDR4 contract prices jumping up to 50 percent in Q1 2026, and the rally has not stopped.
The memory market is doing something it has not done in a decade, and it is doing it for a reason that will not resolve quickly. TrendForce's January 2026 note recorded DDR4 contract prices rising as much as 50 percent quarter-on-quarter, with conventional DRAM forecast to climb 55 to 60 percent in the same quarter. By its July 2026 press release, the pace had moderated — conventional DRAM contract prices forecast at 13 to 18 percent quarter-on-quarter for 3Q26 — but "moderated" here means the third consecutive quarter of double-digit increases, driven by AI server demand pulling wafer capacity away from everything else.
For most PC buyers that is an annoyance. For someone building a local-LLM box it is structural, because the entire hobby is a memory-capacity problem wearing a compute costume. You are not buying FLOPS; you are buying somewhere to put weights. And DDR4 — the memory every AM4 build in this guide uses — is the part of the market under the most pressure, because it sits on shrinking production allocation while the fabs chase HBM.
This guide names five parts that make a working local-LLM machine, ranked by what they solve rather than by price. The winner is the one that ends the argument: a 12 GB graphics card.
Step 0: which tier do you actually need?
Pick your model size before you pick a part. The mapping is tighter than most buying guides admit.
| Model class | Example | Q4_K_M weight size | What it needs | Realistic experience |
|---|---|---|---|---|
| 1B | Gemma 3 1B | 0.81 GB | Anything, including a Pi | 5 tok/s on an SBC, 32 tok/s on a desktop CPU |
| 4B-8B | Llama 3.1 8B | 4.92 GB | 8 GB VRAM, or 16 GB system RAM | 51 tok/s on a 12 GB GPU; 4.8 tok/s CPU-only |
| 14B | Qwen2.5 14B | 8.99 GB | 12 GB VRAM | 26.6 tok/s on a 12 GB GPU; ~4 tok/s CPU-only |
| 30B-A3B (sparse MoE) | Qwen3 30B-A3B | 18.63 GB | 12 GB VRAM plus offload, or 32 GB RAM | Usable with partial offload |
| 30B+ dense | Qwen2.5 32B | 19.85 GB | 24 GB VRAM | Out of scope for this budget |
The 8B and 14B GPU figures come from LocalScore's RTX 3060 submission; the CPU-only figures from its Ryzen 5 5600X and Ryzen 7 5800X submissions. Weight sizes are published artifact sizes from the bartowski GGUF repositories for Qwen2.5 14B, Llama 3.1 8B, Qwen3 30B-A3B and Qwen2.5 32B.
Read the 14B row twice. It is the tier that defines this budget class: 8.99 GB of weights plus a KV cache fits in 12 GB and does not fit in 8 GB. That single fact is why every "cheapest GPU for local LLM" answer converges on the same card.
The picks at a glance
| Pick | Best for | Key spec | Price range | Verdict |
|---|---|---|---|---|
| MSI RTX 3060 12GB | Best overall | 12 GB GDDR6, 170 W | ~$480 | The VRAM tier that matters, at the lowest price it exists |
| ZOTAC RTX 3060 Twin Edge OC 12GB | Best value / availability | 12 GB GDDR6, 192-bit | ~$500 | Same silicon, buy whichever is in stock |
| AMD Ryzen 7 5800X | Best for CPU-offload builds | 8C/16T, 105 W | ~$259 | Cores only help prefill and hybrid offload |
| AMD Ryzen 5 5600X | Best performance per dollar | 6C/12T, 65 W, cooler included | ~$174 | The platform anchor that leaves budget for the GPU |
| AMD Ryzen 5 5600G | Budget pick | 6C/12T with Radeon graphics | ~$200 | No discrete GPU needed for 1B-4B models |
Prices are SpecPicks catalog figures captured 23 September 2026 and may vary; check the live listing before ordering.
Top picks
#1: MSI Gaming GeForce RTX 3060 12GB — Best Overall
Verdict: The cheapest way to cross the 12 GB VRAM line, and the line is what matters. ~$480, 12 GB GDDR6, 170 W.
Per NVIDIA's RTX 3060 family page, this card carries 12 GB of GDDR6 on a 192-bit bus, 3,584 CUDA cores at a 1.78 GHz boost, 170 W of graphics card power, a single 8-pin connector, and a 550 W recommended system supply. Retail listings specify 15 Gbps memory, which on a 192-bit bus works out to 360 GB/s — roughly seven times the 51.2 GB/s a dual-channel DDR4-3200 desktop can manage, and bandwidth is what token generation is starved of.
The measured result: 26.6 tok/s generation on Qwen2.5 14B Q4_K_M with 759 tok/s prefill and a 1.92-second time to first token, per LocalScore #43. The same submission records 51.3 tok/s on Llama 3.1 8B and 185 tok/s on Llama 3.2 1B. Hardware Corner's RTX 3060 12GB tables corroborate the class of result on Qwen3 14B Q4_K: 31.2 tok/s generation and 972.6 tok/s prefill at 4K context, 22.7 and 678.2 at 16K.
Pros
- 12 GB holds 14B-class weights at Q4_K_M with room for a working KV cache
- 170 W and a single 8-pin connector fit inside almost any existing build
- CUDA support means every local-LLM tool works on day one, with no backend hunting
- Six times the CPU-only generation speed, and roughly 29 times better time to first token
Cons
- 192-bit GDDR6 is slow by 2026 standards; a 24 GB card is meaningfully faster as well as larger
- 12 GB runs out at long context on 14B, and at any context above ~14B dense
- Prices on Ampere cards have drifted up, not down, as memory costs rose
Check the live price and stock — prices on this card have been moving with the memory market. Full benchmark data is on the RTX 3060 12GB page.
#2: ZOTAC Gaming GeForce RTX 3060 Twin Edge OC 12GB — Best Value
Verdict: Identical silicon, two-fan cooler, buy whichever of the two is cheaper on the day. ~$500, 12 GB GDDR6, 192-bit.
There is no performance argument to make between AIB partner cards at this tier — same GA106 die, same 12 GB, same 192-bit bus, same 15 Gbps memory. What varies is cooler acoustics, board length and street price, and street price is the only one that should decide the purchase. Every benchmark figure quoted for the MSI card applies here.
Pros
- Two-fan cooler in a shorter board than most triple-fan 3060s, which matters in mini-ITX
- Factory OC on the same silicon
- Genuine alternative when the MSI card is out of stock or over street price
Cons
- Frequently priced above the MSI card despite being the "value" pick
- Same 12 GB ceiling and same 192-bit bandwidth limit
See current pricing. Prices may vary.
#3: AMD Ryzen 7 5800X — Best for CPU-Offload Builds
Verdict: Eight cores, for the specific case where layers live in system RAM. ~$259, 105 W, no cooler included.
Per AMD's specification page, the 5800X is 8 cores and 16 threads with 32 MB of L3, a 105 W default TDP, DDR4 up to 3200 MT/s, and no bundled thermal solution.
Here is the honest case for it, and it is narrower than most guides suggest. Extra cores do almost nothing for token generation, which is memory-bandwidth-bound: three public LocalScore 5800X submissions on Qwen2.5 14B Q4_K_M read 3.5, 4.0 and 4.3 tok/s, against 4.2 for a 6-core 5600X. What cores do scale is prompt processing. In llamafile's CPU benchmark thread, moving from a 6-core 5600X to a 16-core 5950X lifted Mistral 7B prefill 1.77x to 2.01x while generation moved only 1.11x to 1.14x.
So buy the 5800X if you run sparse MoE models that spill out of 12 GB, if you push long documents through a hybrid GPU-plus-CPU setup, or if the same box does compile and encode work between inference jobs. Do not buy it expecting faster chat.
Pros
- 33 percent more cores than the 5600X, which shows up in prefill and hybrid offload
- Same AM4 socket, so it is a drop-in upgrade on an existing B550 board
- The right pick for 30B-A3B-class MoE models that need system-RAM offload
Cons
- No cooler in the box, and a 105 W part under a 100 percent duty cycle needs a real tower
- Measured generation on 14B is indistinguishable from the cheaper 5600X
- 40 W more design power for a workload that runs for hours
Check the live price; benchmark data lives on the Ryzen 7 5800X page.
#4: AMD Ryzen 5 5600X — Best Performance per Dollar
Verdict: The platform anchor. ~$174, 65 W, Wraith Stealth included — and every dollar saved goes into the GPU.
AMD lists the 5600X at 6 cores and 12 threads, the same 32 MB of L3 as the 5800X, a 65 W default TDP, DDR4 up to 3200 MT/s, and an included Wraith Stealth cooler. That cooler is not a rounding error: it is $30 to $45 that does not appear on the 5800X's line item.
Measured CPU-only performance is 4.2 tok/s on Qwen2.5 14B Q4_K_M and 32.4 tok/s on Llama 3.2 1B (LocalScore #1018) — the first of those is fine for overnight batch work and painful for chat, which is precisely why this is the CPU you buy when the GPU is doing the inference.
The counter-case: if your box has no GPU and never will, and you routinely push 4K-token prompts through it, the 5800X's prefill advantage is real and the 5600X is the wrong pick.
Pros
- Best tokens-per-dollar of any part here for CPU-only work: 2.41 tok/s per $100 on 14B
- 65 W and a bundled cooler — genuinely cheap to buy and to run
- Leaves roughly $85 of the build budget for VRAM, which is where it belongs
Cons
- Six cores limits prefill throughput on long prompts
- Same dual-channel DDR4 bandwidth ceiling as every other AM4 part
See current pricing · Ryzen 5 5600X benchmarks
#5: AMD Ryzen 5 5600G — Budget Pick
Verdict: Six cores with Radeon graphics on board, so a 1B-4B assistant needs no discrete card at all. ~$200.
Two public LocalScore submissions put the 5600G at 43.9 and 32.8 tok/s generation on Llama 3.2 1B Q4_K_M — competitive with the 5600X's 32.4 on the same model. But the 5600G's value here is not speed, it is the absence of a second purchase. With integrated graphics it boots and runs headless without a GPU in the slot, which makes it the cheapest complete AM4 machine for small-model work — classification, routing, short summarization, structured extraction. It is also the sensible host if your plan is to add a 12 GB card in six months, because the iGPU keeps the box useful in the meantime.
Where it stops: at 8B and above, CPU-only decode drops into the low single digits and interactive use stops being interactive. It is not a substitute for VRAM.
Pros
- No discrete GPU required, which removes the most expensive line item entirely
- Six Zen 3 cores handle 1B-4B models at usable speeds
- Low idle power for an always-on assistant box
Cons
- Only one public submission clears 40 tok/s even at 1B; the Windows run measured 32.8 with an 8.14-second time to first token
- iGPU inference paths are far less mature than CUDA; expect to fight your backend
- Nothing about it changes the 8B-and-above story
See current pricing · Ryzen 5 5600G benchmarks
What to look for in a local LLM build
VRAM ceiling, before anything else
Every other spec is a tiebreaker. A 12 GB card runs models an 8 GB card cannot load at all, and no amount of extra compute on the smaller card closes that gap. Work out the largest model you actually intend to run, find its Q4_K_M file size, add 2 to 3 GB for a working KV cache, and buy the next VRAM tier up from that number.
Be specific about the SKU, too. The RTX 3060 shipped in both 8 GB and 12 GB versions with different bus widths — 128-bit on the 8 GB card against 192-bit on the 12 GB — and the model number alone does not tell you which you are buying. Check for "12G" or "12GB" in the product title before ordering.
Dual-channel system RAM, and buy it now
Two matched sticks, not one. Single-channel operation roughly halves your memory bandwidth, which halves CPU-only generation and slows every hybrid-offload configuration. 32 GB is the right target for a machine that will offload layers to system RAM.
On timing: DDR4 is a mature product on shrinking allocation. TrendForce's reporting has Samsung, SK hynix and Micron mapping their legacy-DRAM exit while contract prices climb, which means waiting exposes you to supply-driven price moves with no performance upside from a newer part. If your build is AM4, buy the kit with the CPU.
Model-storage speed
Weights load once and then live in page cache, so storage does not affect steady-state throughput. It affects cold-start time and, on a 24/7 box, reliability. A Kingston A400 960GB is the cheap fix — enough capacity for a dozen quantized models and none of the wear-out failure modes of removable media.
PSU headroom
NVIDIA specifies 550 W for an RTX 3060 system. What deserves more attention than the wattage number is the age of the unit: inference is a sustained load, not a bursty one, and a decade-old supply that survived light gaming can drift out of spec under hours of continuous draw.
Cooling for a 100 percent duty cycle
Both the GPU and the CPU will sit at full load for hours at a time. Case airflow that was adequate for gaming often is not adequate here, and the 5800X in particular ships with no cooler at all.
Common pitfalls
- Buying an 8 GB card because it says 3060. The 8 GB variant is a 128-bit part and cannot hold a 14B model. Read the title.
- Buying one 32 GB stick instead of two 16 GB sticks. Single-channel halves your bandwidth on exactly the workload you bought the machine for.
- Leaving XMP/DOCP off. A DDR4-3600 kit runs at 2133 MT/s until you enable its profile. That is a bigger swing than most CPU upgrades.
- Buying CPU cores to make chat faster. Cores scale prefill, not generation. The measured 5600X-vs-5800X gap on 14B generation is inside run-to-run noise.
- Waiting for memory prices to fall before building. Three consecutive quarters of double-digit contract increases is not a dip.
When NOT to buy any of this
If you only ever run 1B-class models for classification or routing, none of these parts are necessary — a single-board computer draws a few watts and does that job. If you need 30B-plus dense models, none of these parts are sufficient either; you are shopping for 24 GB cards, and this guide's budget does not reach. And if you are content with hosted APIs, the honest arithmetic is that $480 buys a very large number of API calls.
Recommended build
Start with the MSI RTX 3060 12GB and the Ryzen 5 5600X, on a 32 GB dual-channel DDR4-3600 kit, with a Kingston A400 960GB for model storage. That combination runs 14B-class models at 26.6 tok/s with a sub-two-second time to first token, costs roughly $650 in silicon before board and memory, and leaves an upgrade path — the 5600X drops out for a 5800X on the same socket if you later need hybrid offload for MoE models.
If the GPU is out of budget this quarter, buy the Ryzen 5 5600G instead, run 1B-4B models on it, and add the card when you can. What you should not do is buy an 8 GB card to save a hundred dollars. The 12 GB line is the entire point of this purchase.
Sources
- LocalScore — NVIDIA GeForce RTX 3060 submission #43
- LocalScore — Ryzen 5 5600X submission #1018
- LocalScore — Ryzen 7 5800X submission #846
- TrendForce — DDR4 leads legacy memory rally, Q1 prices up to 50%
- TrendForce — AI server demand supports memory prices in 3Q26
- NVIDIA — GeForce RTX 3060 family specifications
Related guides
- Best budget GPU for local LLM
- Best CPU for local LLM inference: 5800X vs 5700X vs 5600G
- Which LLMs fit in an RTX 3060 12GB?
- RTX 3060 12GB vs Ryzen 7 5800X on Qwen3 14B
- RTX 3060 12GB benchmark data
Frequently asked questions
Is 12GB of VRAM still enough for local LLMs in 2026?
For the models most people actually run, yes. A 12GB card holds 8B-class weights at Q8_0 with room for context, or 14B-class weights at Q4_K_M with a modest KV cache, and it runs current sparse MoE models with partial offload. Where 12GB stops being enough is long context at higher quantization and dense models above roughly 14B — at that point you are choosing between a 24GB card and CPU offload, not between 12GB variants.
Should I buy DDR4 system memory now or wait?
If your build is AM4, buying now is the lower-risk choice: DDR4 is a mature product on shrinking production allocation, so waiting exposes you to supply-driven price moves without any performance upside from a newer part. Buy a matched dual-channel kit rather than two single sticks, and size it at 32GB — that is the capacity that lets you offload layers to system memory when a model does not fit in VRAM.
Do I need a new power supply for an RTX 3060 12GB?
Usually not. The card's board power sits at 170 W and it uses a single 8-pin connector, so a competent 550W-650W unit handles it alongside a 65W or 105W AM4 CPU. What deserves attention is the age of the supply — inference is a sustained load rather than a bursty one, and a decade-old unit that survived light gaming can drift out of spec under hours of continuous draw. Check the rail age before the wattage number.
Can I skip the GPU entirely with the Ryzen 5 5600G?
For small models, yes. The 5600G's integrated graphics and six cores handle 1B to 4B-class models at usable speeds for classification, routing and short summarization, and it draws far less power than a discrete card at idle. It is not a substitute for a 12GB GPU at 8B and above, where CPU-only decode drops into low single-digit tokens per second and interactive chat stops feeling interactive.
Is a used 12GB card a better buy than a new one right now?
It can be, with two caveats. Inference is a sustained thermal load, so a card that spent two years mining or gaming at high temperatures has less headroom left than its price suggests — ask for the fan and memory-junction history. And a used purchase carries no warranty against the exact failure mode that matters, VRAM degradation. If the discount is under about 25 percent against new street price, the new card is the better risk-adjusted buy.
Citations and sources
- LocalScore — NVIDIA GeForce RTX 3060 submission #43 (accessed 23 September 2026)
- LocalScore — Ryzen 5 5600X submission #1018 (accessed 23 September 2026)
- LocalScore — Ryzen 7 5800X submission #846 (accessed 23 September 2026)
- LocalScore — Ryzen 7 5800X submission #762 (accessed 23 September 2026)
- LocalScore — Ryzen 5 5600G submission #1278 (accessed 23 September 2026)
- TrendForce — "DDR4 Reportedly Leads Legacy Memory Rally with Q1 Prices Up to 50%", 19 January 2026 (accessed 23 September 2026)
- TrendForce — "AI Server Demand Continues to Support Memory Prices in 3Q26", 3 July 2026 (accessed 23 September 2026)
- TrendForce — "Out with the Old: Memory Giants Map Their 2025-26 Exit Strategy amid Supply Crunch" (accessed 23 September 2026)
- NVIDIA — GeForce RTX 3060 family specifications (accessed 23 September 2026)
- AMD — Ryzen 7 5800X product specifications (accessed 23 September 2026)
- AMD — Ryzen 5 5600X product specifications (accessed 23 September 2026)
- mozilla-ai/llamafile — "Lots of CPU benchmarks", Discussion #450 (accessed 23 September 2026)
- Hardware Corner — RTX 3060 12GB LLM benchmarks (accessed 23 September 2026)
- bartowski — Qwen2.5-14B-Instruct-GGUF file sizes (accessed 23 September 2026)
- bartowski — Meta-Llama-3.1-8B-Instruct-GGUF file sizes (accessed 23 September 2026)
- bartowski — Qwen3-30B-A3B GGUF file sizes (accessed 23 September 2026)
- bartowski — Qwen2.5-32B-Instruct-GGUF file sizes (accessed 23 September 2026)
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.
— Mike Perry · Last verified 23 September 2026
