Step 0: find your bottleneck before you buy a CPU
Run this before you spend anything. It takes ten minutes with the llama-bench tool that ships with llama.cpp (README):
- Download a Qwen3 8B Q4_K_M GGUF (5.03 GB in the official Qwen repo).
- Run
llama-bench -m Qwen3-8B-Q4_K_M.gguf -t 1,2,4,6,12 -p 512 -n 128. - Read the
tg128column (generation) and thepp512column (prefill) at each thread count.
How to read the results:
- Generation flattens by 4 to 6 threads. You're bandwidth-bound. A 12-core CPU won't raise that number, but faster RAM will.
- Prefill keeps climbing up to 6 threads. You're compute-bound on prompts. More cores or wider SIMD will help, and a 3900X would too.
- Both are flat and low. Check that you're running two DIMMs in the A2/B2 slots. A single stick halves your bandwidth, and that's the most common self-inflicted wound on these builds.
We expect the 2600 to show the first two patterns together, because that's what every published dual-channel DDR4 run shows. The llama.cpp developer Johannes Gäßler measured it directly on a Ryzen 7 3700X with DDR4-3200: "just 5 threads are enough to fully utilize the memory bandwidth provided by dual channel memory," and performance dropped once threads exceeded physical cores.
Key Takeaways
- Generation is a memory-bandwidth problem. One 3900X ran 6.32, 7.55 and 8.45 tok/s at DDR4-2666, 3200 and 3600 on the same Mistral 7B Q6_K file. That's a 34% swing from RAM alone (llamafile #450).
- Extra threads stop helping early. On an 8-core Ryzen 9 7940HS, going from 4 to 8 threads lifted 70B generation only from 1.30 to 1.37 tok/s (+5%), while prefill rose nearly 80%.
- Prefill is where the 3900X earns its price. We estimate about 60 tok/s for Qwen3 8B against about 20 tok/s on the 2600. That's roughly 3× faster time-to-first-token.
- Estimated Qwen3 8B Q4_K_M generation: about 6.8–8.2 tok/s on a 2600 at DDR4-2933, and about 8.9 tok/s on a 3900X at DDR4-3200. That's a 10–30% gain, and most of it comes from memory speed.
- Used prices on eBay (September 18, 2026): about $50 for the 2600 and about $187 for the 3900X. A used Ryzen 7 5800X is about $190, the same money.
Spec delta: Zen+ 6-core vs Zen 2 12-core
| Part | Cores / threads | L3 cache | Memory support | Street price band (used, eBay US, 2026-09-18) |
|---|---|---|---|---|
| AMD Ryzen 5 2600 (Zen+, 12 nm) | 6 / 12, 3.4–3.9 GHz | 16 MB (2 × 8 MB CCX) | DDR4-2933, 2 channels | $40–65, median ~$50 |
| AMD Ryzen 9 3900X (Zen 2, 7 nm) | 12 / 24, 3.8–4.6 GHz | 64 MB (4 × 16 MB CCX) | DDR4-3200, 2 channels | $150–250, median ~$187 |
| AMD Ryzen 7 5800X (Zen 3, 7 nm) | 8 / 16, up to 4.7 GHz | 32 MB (one CCX) | DDR4-3200, 2 channels | $178–200, median ~$190 |
| AMD Ryzen 5 5600G (Zen 3 APU) | 6 / 12, up to 4.4 GHz | 16 MB | DDR4-3200, 2 channels | $120–200, median ~$136 |
Sources: AMD's spec pages for the Ryzen 5 2600, Ryzen 9 3900X, Ryzen 7 5800X and Ryzen 5 5600G. CCX layout is from Wikipedia's Zen 2 article. Price bands are the interquartile range of used fixed-price listings we pulled from eBay's Browse API; they exclude bundles and lots.
Two rows matter more than the core count. First, the AMD Ryzen 5 2600 is rated for DDR4-2933 and the AMD Ryzen 9 3900X for DDR4-3200. Theoretical dual-channel bandwidth is 46.9 GB/s against 51.2 GB/s (DDR4 transfers 8 bytes per channel per transfer, per Wikipedia's DDR4 table). Second, Zen 2 doubled the floating-point execution width from 128-bit to 256-bit, which "allows the FPU to perform single-cycle AVX2 calculations" (Wikipedia, Zen 2). llama.cpp's CPU matmul kernels are AVX2, so that's a per-clock prefill advantage before core count even enters the picture.
A note on buying new: Amazon's current new-stock listing for the 2600 is priced far above the used market. Buy this chip used.
Qwen3 8B on CPU: measured throughput
No one has published a Qwen3 8B CPU-only run on either chip. The strongest public data is the llamafile thread. It has a 3900X memory-speed sweep on Mistral 7B, and a Zen+ laptop chip (the 4-core Ryzen 5 3550H, a Picasso part per Wikipedia's Ryzen list) on DDR4-2400. The measured rows:
| CPU (arch) | RAM | Model | Prefill pp512 (tok/s) | Generation tg16 (tok/s) | Effective bandwidth |
|---|---|---|---|---|---|
| Ryzen 9 3900X (Zen 2, 12c) | DDR4-2666 | Mistral 7B Q6_K, 5.94 GB | 68.65 | 6.32 | 37.5 GB/s (88% of peak) |
| Ryzen 9 3900X (Zen 2, 12c) | DDR4-3200 | Mistral 7B Q6_K | 69.43 | 7.55 | 44.8 GB/s (88%) |
| Ryzen 9 3900X (Zen 2, 12c) | DDR4-3600 | Mistral 7B Q6_K | 69.26 | 8.45 | 50.2 GB/s (87%) |
| Ryzen 5 3550H (Zen+, 4c laptop) | DDR4-2400 | Mistral 7B Q4_K_M, 4.37 GB | 11.75 | 6.44 | 28.1 GB/s (73%) |
| Ryzen 5 3550H (Zen+, 4c laptop) | DDR4-2400 | Mistral 7B Q6_K | 11.52 | 4.69 | 27.9 GB/s (73%) |
| Ryzen 5 5600X (Zen 3, 6c) | not stated | Mistral 7B Q6_K | 62.33 | 7.51 | — |
All six rows come from llamafile discussion #450. We derived effective bandwidth as model size × generation tok/s, which works because each generated token reads the whole file once.
Look at the 3900X rows. RAM speed moved generation 34% while prefill stayed flat at 69 tok/s. That's the whole thesis in three lines. Then look at the 3550H: four Zen+ cores already pull a steady 28 GB/s on two different quantizations. Even a small Zen+ part saturates what its memory controller delivers.
Applied to Qwen3 8B Q4_K_M (5.03 GB), these are our estimates, not measurements:
| CPU | RAM | Est. prefill (tok/s) | Est. generation (tok/s) | Est. time to first token, 2,000-token prompt |
|---|---|---|---|---|
| Ryzen 5 2600 | DDR4-2933 | ~20 | 6.8–8.2 | ~100 s |
| Ryzen 5 2600 | DDR4-3200 (outside spec, often works) | ~20 | 7.4–9.0 | ~100 s |
| Ryzen 9 3900X | DDR4-3200 | ~60 | ~8.9 | ~33 s |
| Ryzen 9 3900X | DDR4-3600 | ~60 | ~10.0 | ~33 s |
Here's how we got the numbers:
- 3900X generation: its measured 44.8 and 50.2 GB/s divided by 5.03 GB.
- 2600 generation: a range. The low end uses the 73% efficiency Zen+ showed on the 3550H, and the high end uses the 88% Zen 2 showed.
- 2600 prefill: scaled from the 3550H by 1.5× for cores and about 1.2× for desktop clocks.
- 3900X prefill: its measured 69 tok/s, scaled 0.94× for Qwen3 8B's larger matmul parameter count.
Treat all of these as ±15%. Also note that llamafile ships its own CPU matmul kernels, so stock llama.cpp prefill will differ.
Quantization matrix for CPU-only hosts
Qwen3 8B has 36 layers and 8 KV heads with a head size of 128 (config.json). That puts the FP16 KV cache at 144 KiB per token. File sizes are from the Qwen and Unsloth GGUF repos. The tok/s bands are bandwidth math (effective GB/s ÷ file size) and are estimates.
| Quant | File size | System RAM needed (8k context) | Est. gen tok/s, 2600 @2933 | Est. gen tok/s, 3900X @3200 | Quality notes |
|---|---|---|---|---|---|
| Q2_K | 3.28 GB | ~5.5 GB | 10–13 | ~13.7 | Visible degradation; last resort |
| Q3_K_M | 4.12 GB | ~6.5 GB | 8–10 | ~10.9 | Usable for chat, weaker on code |
| Q4_K_M | 5.03 GB | ~7 GB | 6.8–8.2 | ~8.9 | The default sweet spot |
| Q5_K_M | 5.85 GB | ~8 GB | 5.8–7.0 | ~7.7 | Small quality gain, ~15% slower |
| Q6_K | 6.73 GB | ~9 GB | 5.1–6.1 | ~6.7 | Near-Q8 quality |
| Q8_0 | 8.71 GB | ~11 GB | 3.9–4.7 | ~5.1 | Close to lossless |
| BF16 | 16.39 GB | ~19 GB | 2.1–2.5 | ~2.7 | Reference only; pointless on DDR4 |
Two caveats apply. First, low-bit K-quants spend more CPU time dequantizing each byte, so on six Zen+ cores Q2_K and Q3_K_M may fall short of their bandwidth ceiling. Second, the RAM column adds about 1.1 GiB of KV cache at 8k context plus roughly 1 GB for the runtime. It doesn't include your OS.
Why prefill scales with cores and generation does not
During generation, each new token needs one pass through every weight in the model. For Q4_K_M that's 5.03 GB. Dual-channel DDR4-3200 peaks at 51.2 GB/s, so the hard ceiling is about 10 tok/s no matter how many cores are waiting. The 3900X reached 88% of peak with 12 cores. The 3550H reached 73% with four.
The 7940HS thread sweep in the same llamafile thread shows the shape. Going from 4 to 8 threads lifted 70B generation from 1.30 to 1.37 tok/s, and 12 threads added nothing. Prefill on that run went from 2.70 to 4.85 tok/s between 4 and 8 threads, then dipped slightly at 12 once simultaneous multithreading kicked in.
Prefill is different because it batches hundreds of prompt tokens against each weight it loads. Every byte fetched from RAM gets reused many times, so the limit becomes arithmetic throughput. That's where the 3900X has three advantages stacked: twice the cores, twice the AVX2 width per clock, and about 700 MHz more boost. Our ~3× prefill estimate is conservative compared with the raw math.
Qwen3 adds a twist. Its default thinking mode writes a <think> block before answering. That's generation-bound work, often several hundred tokens, and neither CPU speeds it up much. The Qwen3 model card documents a /no_think switch. On a CPU host, use it for anything that isn't reasoning-heavy.
What L3 cache actually changes
The 3900X's 64 MB of L3 looks like a 4× win over the 2600's 16 MB. In practice it's split into four 16 MB slices, one per three-core CCX, and a thread only sees its own slice quickly. The 2600 has two 8 MB slices.
Neither chip can hold a 5 GB weight file in cache, so cache doesn't lift generation. The larger slices help prefill a little, since activation tiles stay resident longer, and they help when other services share the box.
If cache really mattered for this workload, the 5600X in the table above wouldn't come in at 7.51 tok/s, right next to the 3900X's 7.55 at DDR4-3200, with half the total L3 (its RAM speed isn't stated, so treat that as a loose comparison). For a deeper look at the cache question, see our 3D V-Cache CPU-inference analysis.
Context length is the other cliff
At FP16 the KV cache grows by 144 KiB per token:
| Context | KV cache | Total RAM with Q4_K_M weights | Est. generation slowdown vs 2k context |
|---|---|---|---|
| 4k | 0.56 GiB | ~6.5 GB | ~5% |
| 8k | 1.13 GiB | ~7.2 GB | ~10–20% |
| 32k | 4.5 GiB | ~11 GB | ~45–50% |
The slowdown column follows from the same bandwidth logic. Each generated token also reads the whole KV cache, so at 32k you stream about 9.9 GB per token instead of 5 GB, and throughput roughly halves on either CPU. A 16 GB box handles 8k comfortably. For 32k contexts alongside a browser and a desktop, plan on 32 GB.
The 3900X is hit just as hard at long context, and its prefill advantage matters more there. Processing a 30,000-token document would take about 25 minutes on a 2600 against about 8 on a 3900X, by our estimates.
The parts most readers should buy instead
The AMD Ryzen 7 5800X costs the same as a used 3900X on eBay (median about $190). It has eight Zen 3 cores in one 32 MB CCX, 256-bit AVX2, and DDR4-3200 support. The measured six-core Zen 3 5600X already reached about 90% of the 3900X's prefill (62.33 against about 69 tok/s), so eight Zen 3 cores should match or beat it. We estimate 70–75 tok/s on Qwen3 8B. Generation is the same bandwidth-capped ~9 tok/s. The 5800X is also a far better gaming chip if the box does double duty. Our Ryzen 5 2600 vs 5800X upgrade guide covers that side.
The AMD Ryzen 5 5600G wins if your 2600 box has no graphics card. The 2600 has no integrated graphics, and the 5600G's Vega iGPU drives a monitor without a PCIe card. It also brings Zen 3 cores and DDR4-3200 at a median used price of about $136. Our 5600G CPU-inference guide has the details.
Before buying any AM4 upgrade, check that your B450/X470 board has a BIOS that supports Zen 2 or Zen 3. As a 2600 owner, you can flash it with the old chip still installed.
Perf-per-dollar and perf-per-watt
These use estimated Qwen3 8B Q4_K_M throughput at DDR4-3200 (the 2600 at the midpoint of its range) and median used eBay prices. The per-watt figures divide by rated TDP, not measured draw. Wikipedia notes the 3900X can exceed 145 W under load.
| CPU | Used median | Est. gen tok/s | Est. prefill tok/s | Gen tok/s per $100 | Prefill tok/s per $100 | Gen tok/s per TDP watt |
|---|---|---|---|---|---|---|
| Ryzen 5 2600 | $50 | ~8.2 | ~20 | 16.4 | 40 | 0.13 (65 W) |
| Ryzen 5 5600G | $136 | ~8.5 | ~50 | 6.3 | 37 | 0.13 (65 W) |
| Ryzen 9 3900X | $187 | ~8.9 | ~60 | 4.8 | 32 | 0.08 (105 W) |
| Ryzen 7 5800X | $190 | ~8.9 | ~72 | 4.7 | 38 | 0.08 (105 W) |
The 2600 you already own is the value leader because its price is sunk. Every upgrade buys prefill, and none of them buys meaningful generation.
Verdict matrix
- Keep the Ryzen 5 2600 if your prompts are short (chat, quick questions, fewer than 500 tokens) and you can put a DDR4-3200 kit in two channels. You'll get about 7–8 tok/s generation, which feels like steady typing. Spend the $140 on RAM or save it toward a GPU.
- Get the Ryzen 9 3900X if you feed the model long inputs: RAG over documents, code-repo context, or summarizing pages. Cutting time-to-first-token by about 3× is the real upgrade. It's also the pick if the box runs other heavily threaded services (transcoding, CI builds) alongside inference.
- Get neither and save for a 12 GB GPU if you want interactive speed. A GPU-resident 8B model on an RTX 3060 12GB runs several times faster than any AM4 CPU. Our RTX 3060 model-fit guide lists what fits.
The recommended pick
For Qwen3 8B, keep the Ryzen 5 2600 and put about $60 into a matched 2×16 GB DDR4-3200 kit. You'll recover most of the generation gap a 3900X would give you (about 7.4–9 against 8.9 tok/s) for a third of the price. The condition that flips it is prompt length. If more than a third of your requests carry 1,000+ tokens of context, the prefill wait dominates the experience, and a CPU upgrade is worth it. At that point, buy the Ryzen 7 5800X over the 3900X: it costs the same used, is faster per core, and is a better gaming chip.
Bottom line
Twelve cores don't make an 8B model talk faster on dual-channel DDR4. The memory bus decides that, and published 3900X runs show RAM speed moving generation 34% while its cores sat idle. What twelve cores buy is roughly three times faster prompt processing, and that only matters if your prompts are long. Measure your own thread scaling with llama-bench before spending anything. Fix your RAM first. Upgrade the CPU only for prefill, and when you do, prefer Zen 3.
Related guides
- Ryzen 9 3900X vs Ryzen 5 5600G for CPU-only local LLM inference
- How many CPU cores does a local-LLM rig actually need?
- Dual-channel RAM for local LLM inference
- Best CPU for a local-LLM homelab under $300
- Local LLM CPU-only on a Ryzen 7 5800X
- Ryzen 5 2600 vs Ryzen 7 5800X for CPU-only Qwen3 30B-A3B: the MoE version of this question
Live price comparison
For side-by-side current pricing and buy buttons, see Ryzen 5 2600 vs Ryzen 9 3900X head-to-head. You can also check per-chip results on the Ryzen 9 3900X benchmark page and the Ryzen 5 2600 benchmark page.
Citations and sources
- llamafile discussion #450, "Lots of CPU benchmarks" — 3900X memory sweep, 3550H, 5600X and 7940HS thread rows (accessed 2026-09-18)
- Johannes Gäßler, llama.cpp Performance Testing — thread saturation on dual-channel memory (accessed 2026-09-18)
- AMD Ryzen 5 2600 specifications (accessed 2026-09-18)
- AMD Ryzen 9 3900X specifications (accessed 2026-09-18)
- AMD Ryzen 7 5800X specifications (accessed 2026-09-18)
- AMD Ryzen 5 5600G specifications (accessed 2026-09-18)
- Wikipedia: Zen 2 — 256-bit FPU and CCX L3 layout (accessed 2026-09-18)
- Wikipedia: List of AMD Ryzen processors — 3550H is Picasso (Zen+); 3900X load power (accessed 2026-09-18)
- Wikipedia: DDR4 SDRAM — per-channel bandwidth (accessed 2026-09-18)
- Qwen/Qwen3-8B model card, Qwen3-8B-GGUF and unsloth/Qwen3-8B-GGUF — architecture and quant file sizes (accessed 2026-09-18)
- llama.cpp llama-bench README (accessed 2026-09-18)
Editorial synthesis: SpecPicks did not bench-test these CPUs for this article. Measured figures are credited inline to their public sources. Figures marked as estimates are our bandwidth- and scaling-based projections, and the method is shown next to each table.
