Introduction: the GPU-less starting point
Plenty of local LLM boxes never get a graphics card. Some are homelab servers that already run Jellyfin or Home Assistant. Some are always-on machines where adding a 170 W graphics card is hard to justify. And some just belong to people who want to try Ollama or llama.cpp before they spend $250 or more on a 12GB card. If you are in any of those groups, the CPU and the RAM are the whole inference stack. The two AM4 chips people most often have lying around, or find cheapest on the used market, are the Ryzen 9 3900X and the Ryzen 5 5600G.
On paper it looks lopsided. The 3900X has 12 cores and 24 threads, a 64 MB L3 cache and a 105 W TDP per AMD's 3900X spec listing. The 5600G has 6 cores and 12 threads, 16 MB of L3 and a 65 W TDP per AMD's 5600G spec listing. Anyone who has watched a Cinebench score would expect the 3900X to win easily.
"Usable" needs a definition first. As a working rule, generation above about 8 tok/s feels like live chat. Around 3 to 5 tok/s feels like watching someone type. Below 2 tok/s you are running a batch job. Prompt processing, also called prefill, matters separately. It sets how long you wait before the first word appears, and that wait grows with long documents and agent loops.
A note on sourcing: SpecPicks did not run first-party benchmarks for this piece. Every number below comes from a public source linked inline, or from arithmetic on those numbers that is labelled as an estimate. No public benchmark we could find tests the 5600G itself on CPU inference. So we use the Ryzen 5 5600X, which has the same Zen 3 cores, the same core count and the same dual-channel DDR4-3200 rating, as the measured proxy, and we call out where the two chips differ.
Key takeaways
- 12 cores vs 6 cores: 12% faster generation. Mistral 7B Q6_K on DDR4-3600 ran at 8.45 tok/s on the 3900X and 7.51 tok/s on a six-core Zen 3 5600X (llamafile #450).
- 16 Zen 3 cores didn't help either. A Ryzen 9 5950X on DDR4-3600 generated 8.50 tok/s on the same model, within 1% of the 3900X.
- RAM speed moved the 3900X by 34%. On one 3900X, going from DDR4-2666 to DDR4-3600 lifted generation from 6.32 to 8.45 tok/s, while prefill stayed flat at 68.65 to 69.26 tok/s.
- Prefill is where cores pay: +11%. The 3900X processed 512-token prompts at 69.26 tok/s vs 62.33 tok/s for the six-core Zen 3 part.
- Dual-channel DDR4-3200 tops out at 51.2 GB/s, which caps a 4.58 GiB Llama 3.1 8B Q4_K_M file at roughly 10 tok/s of generation on either chip.
- Power: 65 W vs 105 W TDP. The 5600G does about the same generation work at 62% of the 3900X's rated TDP.
Step 0: is your bottleneck cores or memory bandwidth?
To generate each token, a dense transformer reads essentially all of its weights once. For a quantized model, that means every generated token needs about one model file's worth of bytes pulled from RAM into the CPU. The rate at which your memory can deliver those bytes is the real speed limit. llama.cpp developer Johannes Gäßler says so directly: for CPU inference "the most important factor is memory bandwidth," and "the actual CPU doesn't matter much" (llama.cpp performance notes).
The arithmetic is simple. DDR4-3200 is PC4-25600, a 25.6 GB/s module rating per Wikipedia's DDR4 SDRAM article. Two channels double that to 51.2 GB/s. Both the 3900X and the 5600G are officially rated for DDR4-3200 in dual channel, so both have the same 51.2 GB/s theoretical ceiling at stock settings.
Llama 3.1 8B Instruct at Q4_K_M is a 4.58 GiB file, or 4.92 GB, per the model listing in the llamafile benchmark thread. Divide the bandwidth by the model size:
- Theoretical ceiling at DDR4-3200: 51.2 ÷ 4.92 ≈ 10.4 tok/s
- Theoretical ceiling at DDR4-3600: 57.6 ÷ 4.92 ≈ 11.7 tok/s
Real systems land below the ceiling. Checking the 3900X's measured Mistral 7B Q6_K runs (5.53 GiB, or 5.94 GB) against that math shows how far below:
| RAM speed | Theoretical bandwidth | Measured gen tok/s | Effective GB/s | Share of ceiling |
|---|---|---|---|---|
| DDR4-2666 | 42.7 GB/s | 6.32 | 37.5 | 88% |
| DDR4-3200 | 51.2 GB/s | 7.55 | 44.8 | 88% |
| DDR4-3600 | 57.6 GB/s | 8.45 | 50.2 | 87% |
The ratio holds steady at 87 to 88% across all three speeds. That steadiness is the proof: generation is tracking memory speed, not compute. If the 12 cores were the limit, faster RAM would have stopped helping. Instead, each step up in RAM speed moved throughput by almost exactly the bandwidth gain.
Estimate: at 88% efficiency on DDR4-3200, Llama 3.1 8B Q4_K_M lands around 9 tok/s of generation on either chip. That is an estimate from the math above, not a measurement.
Spec-delta table: the two chips side by side
| Chip | Cores / threads | Base / boost clock | L3 cache | TDP | Memory support |
|---|---|---|---|---|---|
| Ryzen 9 3900X (Zen 2) | 12 / 24 | 3.8 / 4.6 GHz | 64 MB | 105 W | DDR4-3200, dual channel, PCIe 4.0 |
| Ryzen 5 5600G (Zen 3 APU) | 6 / 12 | 3.9 / 4.4 GHz | 16 MB | 65 W | DDR4-3200, dual channel, PCIe 3.0 |
| Ryzen 7 5800X (Zen 3) | 8 / 16 | 3.8 / 4.7 GHz | 32 MB | 105 W | DDR4-3200, dual channel, PCIe 4.0 |
| Ryzen 5 2600 (Zen+) | 6 / 12 | 3.4 / 3.9 GHz | 16 MB | 65 W | DDR4-2933, dual channel, PCIe 3.0 |
Sources: AMD's spec listings for the 3900X, 5600G and 5800X, plus Wikipedia's list of AMD Ryzen processors for clock rates, the Ryzen 2000 series' DDR4-2933 rating and platform PCIe versions.
Two rows matter most. The memory column is identical for the two chips being compared, so they share the same generation ceiling. The Ryzen 5 2600 has a lower rating (DDR4-2933, or 46.9 GB/s theoretical), so at stock it starts about 8% behind. The core column is where the 3900X leads, and the next sections show where that lead actually shows up.
Benchmark table: CPU-only throughput by model size
All rows come from user-submitted llamafile-bench runs in llamafile discussion #450. "Prefill" is pp512 (processing a 512-token prompt) and "gen" is tg16 (generating 16 tokens). The runs used different llamafile versions and system builds, so read differences under about 10% as noise.
| Chip (RAM) | Model | Prefill tok/s | Gen tok/s | Source |
|---|---|---|---|---|
| Ryzen 9 3900X (32 GB DDR4-3600) | TinyLlama 1.1B F16 | 278.28 | 23.10 | llamafile #450 |
| Ryzen 9 3900X (32 GB DDR4-3200) | Mistral 7B Q6_K | 69.43 | 7.55 | llamafile #450 |
| Ryzen 9 3900X (32 GB DDR4-3600) | Mistral 7B Q6_K | 69.26 | 8.45 | llamafile #450 |
| Ryzen 5 5600X, 5600G proxy (96 GB DDR4-3600, 4 DIMMs) | Mistral 7B Q6_K | 62.33 | 7.51 | llamafile #450 |
| Ryzen 5 5600X, 5600G proxy (32 GB DDR4-3000) | Mistral 7B Q4_0 | 23.70 | 8.95 | llamafile #450 |
| Ryzen 9 5950X (DDR4-3600) | Mistral 7B Q6_K | 109.37 | 8.50 | llamafile #450 |
| Ryzen 9 5950X (DDR4-3600) | Llama 3.1 8B Q4_K_M | 100.20 | 10.53 | llamafile #450 |
| Ryzen 9 3900X (32 GB DDR4-3200) | Mixtral 8x7B Q4_K_M (~13B active) | 39.22 | 5.58 | llamafile #450 |
| Ryzen 5 5600X, 5600G proxy (96 GB DDR4-3600, 4 DIMMs) | Mixtral 8x7B Q5_K_M (~13B active) | 31.70 | 5.02 | llamafile #450 |
| Ryzen 5 5600X, 5600G proxy (96 GB DDR4-3600, 4 DIMMs) | Qwen2.5 Coder 14B F16 | 14.64 | 1.48 | llamafile #450 |
How to read it:
- Generation clusters by memory, not by core count. The 7B Q6_K row gives 7.51 (6 cores), 8.45 (12 cores) and 8.50 (16 cores), all on DDR4-3600. The six-core system was running four mismatched DIMMs, which usually costs some effective bandwidth, so part of even that 12% gap is probably the memory setup.
- The 5950X Llama 3.1 8B Q4_K_M row confirms the ceiling math. 10.53 tok/s × 4.92 GB is 51.8 GB/s, or 90% of DDR4-3600's 57.6 GB/s. Sixteen fast cores did not break past memory bandwidth.
- The 14B F16 row shows the cliff. A 27.51 GiB file on dual-channel DDR4 gives 1.48 tok/s. A 14B at Q4_K_M is roughly a third of that size, so expect about three times the speed, or around 4 to 5 tok/s (estimate). That is usable for batch jobs but slow for chat.
- Mixture-of-experts models are the CPU-friendly exception. Mixtral 8x7B reads only its active experts per token, so a 26.49 GiB file still generated 5.58 tok/s on the 3900X. That is why MoE models such as gpt-oss 20B are popular on CPU-only boxes.
Does the 5600G's integrated GPU change the answer?
The 5600G is an APU. It has Radeon Graphics with 7 GPU cores per AMD's 5600G listing, and llama.cpp can target it through its Vulkan backend, documented in llama.cpp's build guide. The 3900X has no integrated graphics at all.
The iGPU does not raise the generation ceiling. It has no memory of its own. It borrows system RAM through the same dual-channel DDR4-3200 controller, so it is bound by the same 51.2 GB/s figure worked out above. If 4.92 GB of weights take the CPU cores about 0.1 seconds per token to stream, they take the iGPU about the same.
Where the iGPU can help is prefill, which is compute-bound. Offloading that matrix math to GPU cores leaves the CPU threads free for other homelab work. We could not find a published, sourced 5600G Vulkan prefill measurement, so we are not quoting one. Treat it as something to test on your own box: run llama-bench once with the Vulkan build and once CPU-only, on the same model, and compare the pp512 column. Our Ryzen 5 5600G iGPU local-LLM guide covers setup.
One practical caveat: the iGPU's memory reservation comes out of your system RAM. On a 16 GB system, carving 2 GB for UMA graphics leaves less room for the model and KV cache.
Thread-count tuning: why -t 12 is usually the wrong flag
More threads is not free. Gäßler measured this on a Zen 2 Ryzen 7 3700X with dual-channel DDR4, and found that "just 5 threads are enough to fully utilize the memory bandwidth." He also saw "a noticeable drop in performance when going from 8 to 9 threads," once the thread count exceeded the chip's physical cores (llama.cpp performance notes).
For our two chips, that means:
- Ryzen 5 5600G: start generation at
-t 6, one thread per physical core. SMT threads beyond 6 compete for the same memory stream and add scheduling overhead. - Ryzen 9 3900X: try
-t 6,-t 8and-t 12. Generation will likely plateau by 6 to 8 threads. Prefill can keep scaling toward 12. - Never default to all 24 threads on the 3900X. Going past physical cores was a loss in Gäßler's testing.
There is a topology point here too. Wikipedia's Ryzen listing gives the 3900X a "4 × 3" core configuration: four core complexes of three cores each, spread across two chiplets (CCDs), with memory accessed through a separate I/O die. The 5600G is a single monolithic die with all six cores in one complex. Threads spread across both 3900X chiplets add cross-die hops, which is another reason a smaller, pinned thread count often beats "use everything."
Prefill versus generation: which chip you feel in an agent loop
Look at the two columns separately and the choice gets clearer.
Generation is what you feel in chat, where the prompt is short and the reply is long. Here the chips are close to tied, at 8.45 vs 7.51 tok/s on 7B Q6_K.
Prefill is what you feel with long inputs. An agent loop that re-sends an 8,000-token context on every step, a RAG pipeline that stuffs in retrieved documents, and a code assistant reading a whole file all hit prefill hard. Using the measured Mistral 7B Q6_K prefill rates:
| Prompt length | Ryzen 9 3900X (69.26 tok/s) | Six-core Zen 3 (62.33 tok/s) |
|---|---|---|
| 1,000 tokens | 14 s | 16 s |
| 4,000 tokens | 58 s | 64 s |
| 8,000 tokens | 1 min 56 s | 2 min 8 s |
These wait times are simple division, so read them as estimates. Real prefill slows somewhat as context grows. The 3900X saves you about 12 seconds on an 8K prompt, and the gap is proportionally the same at every length. Neither chip turns an 8K agent context into an interactive experience. That job needs a GPU, or a much wider CPU, like the 5950X's measured 109.37 tok/s prefill.
Perf-per-dollar and perf-per-watt on the used AM4 market
Prices below are current SpecPicks catalog listings as of September 2026. Generation figures are the measured Mistral 7B Q6_K DDR4-3600 numbers, and the 5600G uses its 5600X proxy value.
| Chip | Listed price | Gen tok/s | $ per gen tok/s | Gen tok/s per TDP watt |
|---|---|---|---|---|
| Ryzen 9 3900X | $228.50 | 8.45 | $27.04 | 0.080 |
| Ryzen 5 5600G | $199.99 | 7.51 (proxy) | $26.63 | 0.116 |
| Ryzen 7 5800X | $254.90 | not measured | n/a | n/a |
| Ryzen 5 2600 | $265.00 | not measured | n/a | n/a |
What the table says:
- Dollar for dollar, it's a tie at about $27 per token/second.
- Per watt, the 5600G wins by about 45% against rated TDP. Wikipedia's Ryzen listing also notes the 3900X "may consume over 145 W under load," so the real-world gap under a sustained inference load can be wider. For an always-on box, that difference shows up on your power bill every month.
- The Ryzen 7 5800X is the step-up option, but only for prefill. It has the same DDR4-3200 ceiling, so expect generation within noise of the other two. Its 8 Zen 3 cores should land its prefill between the six-core Zen 3 part and the 5950X. Our 5800X vs 5700X vs 5600G CPU inference comparison goes deeper.
- The Ryzen 5 2600 is the floor, and the listed price is wrong for it. $265 is above its $199 launch price per Wikipedia. It has a lower memory rating and older Zen+ cores. Only consider it used, well under $100, or if it is already in the drawer.
Common pitfalls
- Running one stick of RAM. Single-channel halves bandwidth, which roughly halves generation on either chip. Two matched DIMMs beat one larger DIMM every time.
- Leaving XMP/EXPO off. Many DDR4-3200 kits boot at DDR4-2133 or 2400 without XMP. On the 3900X, 2666 vs 3600 was a 34% generation difference in the benchmark thread.
- Buying the 3900X to speed up chat. Twelve cores barely move generation speed. If your workload is short prompts and long answers, the extra cores sit idle.
- Picking a model that doesn't fit in RAM. Once weights spill into swap, throughput collapses. Budget the model file plus 2 to 4 GB for runtime and KV cache.
- Using F16 weights on CPU. The 14B F16 row shows 1.48 tok/s. Q4_K_M cuts the file size by about two-thirds and roughly triples generation speed.
Which chip should you buy?
Get the Ryzen 9 3900X if…
- your workload is prefill-heavy: RAG over long documents, agent loops re-reading context, or batch summarization
- the box also compiles code, transcodes or runs VMs, where 12 cores pay off beyond inference
- you plan to add a GPU later and want PCIe 4.0 rather than the 5600G's PCIe 3.0
Get the Ryzen 5 5600G if…
- the machine is an always-on homelab server where 65 W TDP and idle behaviour matter
- your workload is chat with short prompts, where generation speed is basically tied
- you want integrated graphics for a display-less build, plus a Vulkan iGPU path to experiment with prefill
Get neither and add a GPU if…
- you want 8B-class models at more than 20 tok/s, or 14B models at conversational speed
- your prompts routinely exceed 4,000 tokens and 60-second waits are unacceptable
Our pick: for a CPU-only local LLM box, the Ryzen 5 5600G. At 7.51 tok/s vs 8.45 tok/s on the same RAM speed, it gives up about 11% of generation speed. In return it costs $28.51 less, is rated for 40 W less, and adds an iGPU. Spend the savings on a fast, matched dual-channel DDR4-3600 kit, which moved generation by 34% in the data above, far more than doubling the cores did.
Bottom line
For CPU-only inference, memory bandwidth decides generation speed, and both chips share the same dual-channel DDR4 ceiling. Expect roughly 9 tok/s on an 8B Q4_K_M model on DDR4-3200 (estimate), and about 7.5 to 8.5 tok/s on a 7B Q6_K on DDR4-3600 (measured). The Ryzen 9 3900X's extra six cores buy about 11% faster prompt processing and little else. Buy the Ryzen 5 5600G at $199.99, spend the difference on fast dual-channel RAM, and put the next real money toward a GPU.
Related guides
- RTX 3060 12GB vs Ryzen 9 3900X for Qwen2.5 14B
- How many CPU cores does a local LLM rig need? 5600G vs 5800X
- Running local LLMs on a Ryzen 5 5600G with no GPU
- Best parts for a CPU-offload local LLM build
- Ryzen 7 5800X CPU inference vs a 12GB GPU
Live price comparison
Check current pricing for both chips on the Ryzen 9 3900X vs Ryzen 5 5600G head-to-head page. For the step-up and floor alternatives, see the Ryzen 7 5800X and Ryzen 5 2600 product pages.
Citations and sources
- mozilla-ai/llamafile Discussion #450, "Lots of CPU benchmarks" (accessed 2026-09-17)
- Johannes Gäßler, llama.cpp Performance Testing (accessed 2026-09-17)
- AMD, Ryzen 9 3900X specifications (accessed 2026-09-17)
- AMD, Ryzen 5 5600G specifications (accessed 2026-09-17)
- AMD, Ryzen 7 5800X specifications (accessed 2026-09-17)
- Wikipedia, List of AMD Ryzen processors (accessed 2026-09-17)
- Wikipedia, DDR4 SDRAM (accessed 2026-09-17)
- ggml-org/llama.cpp, build documentation (Vulkan backend) (accessed 2026-09-17)
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported. Rows labelled "5600G proxy" are measurements of a Ryzen 5 5600X, and figures labelled "estimate" are derived by arithmetic from the cited measurements.
