For a single-user local LLM on the current dense open-weight generation, an AMD Strix Halo APU's 128 GB unified LPDDR5X memory pool beats a 12 GB RTX 3060 on models above 12 GB and loses on models below it. The 3060's ~360 GB/s of dedicated GDDR6 bandwidth still edges the unified LPDDR5X pool on generation throughput for anything that fits. The right choice depends entirely on which weight file you plan to load, not on marketing about "unified memory being the future."
Who this is for
You've watched AMD's Strix Halo launch and you're weighing it against the used or entry-level discrete-GPU rig on your desk. You're a single user — one prompt at a time, no batching, no serving traffic. You want a real number: which one is faster, and where does the crossover live? We synthesize public benchmarks and manufacturer bandwidth specs to lay out the two architectures side by side.
We assume you already know what a KV cache is and what q4_K_M means. If you don't, start with our companion piece — Qwen 3.8 open weights on a 12 GB RTX 3060 — and come back.
Key takeaways
- Strix Halo wins on capacity: 128 GB of shared LPDDR5X means you can load models a discrete 12 GB card cannot even see
- The RTX 3060 12 GB wins on generation throughput for anything that fits its VRAM, thanks to 360 GB/s of dedicated GDDR6
- Crossover is at ~11-12 GB of model size — below that, dGPU wins; above, unified memory is the only game in town without CPU offload
- Power and form factor are the sleeper story: Strix Halo runs a full LLM at 70-120 W in a laptop-class thermal envelope
- Neither architecture is "better" — pick by the weight file, not the acronym
Step 0 — the two architectures at a glance
Discrete GPU (RTX 3060 12 GB). GDDR6 memory sits on the same board as the GPU die, wired over a 192-bit bus at 360 GB/s of theoretical bandwidth per the TechPowerUp spec sheet. The card is fast at what fits and unusable for what doesn't. Anything that doesn't fit in 12 GB spills to the CPU at ~50 GB/s of DDR4 bandwidth — a 7× penalty.
Unified memory (Strix Halo). AMD's Strix Halo APU pairs a Zen 5 CPU cluster with an RDNA 3.5 iGPU on one die, and connects both to a shared LPDDR5X-8000 memory pool documented on AMD's Ryzen AI product page. Total system bandwidth is in the ~200-256 GB/s range depending on channel width and configuration — well above a normal APU, roughly two-thirds of a mid-range discrete card. But every GB of that pool is addressable by the iGPU, not just the 12-16 GB you'd get on a dGPU.
Spec-delta table
| Config | Compute memory | Bandwidth | TDP | Peak $ | Verdict |
|---|---|---|---|---|---|
| RTX 3060 12 GB (dGPU) | 12 GB GDDR6 | 360 GB/s | 170 W | $310 | Fast for models ≤11 GB |
| Strix Halo 128 GB | 128 GB LPDDR5X-8000 | ~200-256 GB/s | 55-120 W | $2000+ (mini-PC) | Fits 30B+ dense at 4-bit |
| Ryzen 7 5800X CPU-only | ≤128 GB DDR4-3200 | ~50 GB/s | 105 W | $180 CPU only | Slow at any scale |
| Reference 24 GB dGPU | 24 GB GDDR6X | 900 GB/s | 350 W | $1000+ used | Fits 13B at high quant |
Note the 24 GB reference row: an older 24 GB card wins on bandwidth but costs 2-4× more than either option and pulls 3× more power than Strix Halo. It exists as a useful third data point, but the real fight is the first two rows.
What can you actually load?
Model file sizes on a q4_K_M quant, using publicly reported measurements for the dense open-weight generation of 2026:
| Model class | q4_K_M weight | KV @ 16K | Total footprint | RTX 3060 12 GB | Strix Halo 128 GB |
|---|---|---|---|---|---|
| 7B dense | ~4.5 GB | ~2 GB | ~7 GB | Fits comfortably | Fits — massive overhead |
| 8B dense | ~5.2 GB | ~2 GB | ~7.5 GB | Fits comfortably | Fits |
| 13B dense | ~8.0 GB | ~2 GB | ~10.5 GB | Fits tight | Fits |
| 30B dense | ~19 GB | ~3 GB | ~22 GB | Does not fit | Fits — 100 GB spare |
| 70B dense | ~42 GB | ~4 GB | ~46 GB | Does not fit | Fits — 80 GB spare |
| 100B MoE (active ~20B) | ~60 GB | ~5 GB | ~65 GB | Does not fit | Fits — tight but real |
| 400B MoE | ~230 GB | ~10 GB | ~240 GB | Does not fit | Does not fit |
The 3060 hits its ceiling at 13B dense; Strix Halo hits its ceiling at ~100B MoE. That's not a marginal difference — it is a two-generation gap in what you can load.
Throughput — the second-order comparison
Once a model fits, generation throughput is a function of memory bandwidth ÷ weight file size, times an efficiency factor (compute overhead, attention math, driver quirks). Approximate numbers using public llama.cpp benchmarks and equivalent iGPU testing on Strix Point / Strix Halo hardware from Phoronix's 2026 compute review:
| Model | RTX 3060 12 GB tok/s | Strix Halo tok/s | Winner |
|---|---|---|---|
| 7B dense q4 | 40-50 | 28-38 | 3060, by ~30% |
| 8B dense q4 | 35-45 | 25-33 | 3060 |
| 13B dense q4 | 22-28 (very tight) | 18-24 | 3060, narrowly |
| 30B dense q4 | N/A (does not fit) | 8-12 | Strix Halo, uncontested |
| 70B dense q4 | N/A | 3-5 | Strix Halo |
| 100B MoE | N/A | 4-8 | Strix Halo |
The pattern is clean: the 3060 wins on tok/s wherever it can play. Strix Halo wins wherever the 3060 forfeits. There is no configuration where Strix Halo is faster than a discrete GPU on a model that fits both.
Prefill vs generation — which one you feel
Prefill (processing your input prompt) is compute-bound. Strix Halo's 40 CU RDNA 3.5 iGPU has roughly 4× the raw compute of a 3060 — the RTX 3060 delivers 12.7 shader TFLOPS per its TechPowerUp entry, while Strix Halo's iGPU lands in the 50+ TFLOPS range at full boost. On very long prompts (16K-32K tokens), the Halo iGPU actually processes prefill faster than the 3060, even though it generates tokens slower.
For chat workflows with short prompts, generation is what you feel — Strix Halo lags. For document-summarization workflows with 32K prompts, prefill dominates — Strix Halo pulls even or ahead once you factor in per-token latency for the small output.
Context length — where unified memory gets a second win
The KV cache on a 30B dense model at 32K context is ~6 GB. On a 3060, that alone doesn't fit alongside the weights. On Strix Halo, you can push context to 100K+ tokens without approaching the memory ceiling.
If your workflow is "paste a whole codebase, ask questions" or "summarize a 200-page PDF," the 3060 can't do it at all on models bigger than 8B. Strix Halo can do it on models up to 70B dense.
The build — parts that pair with each side
The 3060 12 GB pairs with a mid-range Ryzen 7 desktop. See our companion piece for the full parts list. For this article, we'll compare two 2026 builds:
- 3060 12 GB build: MSI GeForce RTX 3060 Ventus 3X 12G + AMD Ryzen 7 5800X + Crucial BX500 1TB SATA SSD for the model library + Noctua NH-U12S. Total: ~$800-900.
- Strix Halo mini-PC: Available as a pre-built box from Framework, Beelink, and Minisforum in 2026 with 64/96/128 GB LPDDR5X. 128 GB SKUs land around $2000-2400 depending on the vendor. No parts swaps needed; the memory is soldered.
The Strix Halo mini-PC pulls 55-120 W at wall for a full LLM under generation, vs 250-350 W for the desktop build with the 3060 spun up. On a laptop battery, that's the difference between three hours of chat and 30 minutes.
Real-world numbers — where each config fits
Three concrete scenarios you might actually build for:
- Chat + code review at 8K context, 8B model. RTX 3060: 35-45 tok/s, snappy. Strix Halo: 28-33 tok/s, slightly slower but fine. Winner on speed: 3060. Winner on cost: 3060 by ~$1500.
- 32K document summarization on a 30B model. RTX 3060: cannot load. Strix Halo: 10 tok/s, works. Winner: Strix Halo, no contest.
- RAG over a 70B model with 16K context. RTX 3060: cannot load. Strix Halo: 4 tok/s. Winner: Strix Halo — but at 4 tok/s you're an API customer, not a local user.
The 3060 is faster where it plays. Strix Halo plays where the 3060 forfeits. There is no scenario where you'd take a Strix Halo mini-PC over a 3060 for the same model. The question is which side of the fit/no-fit line your workload lives on.
Perf-per-dollar and perf-per-watt
For the model class each config runs comfortably:
| Config | Best-fit model | Tok/s | Total system $ | Watts | $/tok/s | Tok/s per W |
|---|---|---|---|---|---|---|
| RTX 3060 desktop | 8B dense q4 | 40 | $850 | 250 | $21.25 | 0.16 |
| Strix Halo mini-PC | 30B dense q4 | 10 | $2200 | 90 | $220.00 | 0.11 |
$/tok/s heavily favors the 3060 — because tok/s on an 8B model is much higher than on a 30B model. But you're comparing different workloads. The right way to read the table: if 30B dense is what you need, Strix Halo's $220/tok/s is the only option in that price band that doesn't involve a used $1500 24 GB card.
Tok/s per watt: the 3060 wins narrowly, but the absolute wattage on Strix Halo is ~3× lower. If you're building a machine to leave running 24/7, the electricity bill favors Halo.
Common pitfalls
- Assuming unified memory is automatically faster. It isn't. LPDDR5X-8000 is fast for laptop RAM but still slower than GDDR6 on a discrete card. Unified memory buys you capacity, not throughput.
- Assuming a 3060 can "kinda run" a 30B model with offload. It can't — the moment you offload, generation drops to CPU-DRAM speed (~5 tok/s), which is worse than running on Strix Halo natively.
- Ignoring the driver stack. As of 2026, ROCm on RDNA 3.5 iGPUs is production-quality for
llama.cpp, but hobby-project fine-tuning frameworks (Unsloth, Axolotl) still assume CUDA. Verify your specific workflow. - Buying Strix Halo for the CPU alone. The Zen 5 cluster inside Strix Halo is good, but the value is in the iGPU + shared memory. For CPU-only work, a bare Ryzen 9 desktop is cheaper and faster.
When NOT to buy each
- Don't buy the 3060 12 GB if: you need to load anything above ~13B parameters at usable throughput, you want long context (32K+) on models above 8B, or you want to fine-tune.
- Don't buy Strix Halo if: your only workload is 7B-13B chat and you already have a desktop, you need CUDA-only workflows, or you can't tolerate 25-30 tok/s (many users can't).
Verdict matrix
Pick RTX 3060 12 GB if… you're building for the dense 7-13B open-weight class, want the fastest chat response for your dollar, and don't need 30B+.
Pick Strix Halo if… you want to run 30B-100B models locally at usable-if-not-fast throughput, care about power/thermal envelope, or need long context on big models.
Pick neither if… you can afford a used 24 GB card ($800-1200 for a 3090). It's faster than either on every model that fits both.
Bottom line
Strix Halo's 128 GB unified pool is the first consumer-adjacent chip that lets a single user run 70B dense models locally without a workstation card. Its throughput on those models is not fast — 3-5 tok/s — but it's real, and the RTX 3060 12 GB simply cannot play in that arena. Below the 12 GB line, the 3060 stays the perf-per-dollar champion. Above it, Strix Halo is the only game in town short of a $1000+ used card. Pair either with a Ryzen 7 5800X or its equivalent, a Crucial BX500 SATA SSD for the model library, and a Noctua NH-U12S for sustained-load thermals, and you have a local-LLM machine that punches above its price.
Related guides
- Qwen 3.8 Open Weights on a 12 GB RTX 3060
- Best GPU Under $400 for Local Llama 8B
- vLLM vs Ollama for a Single-User 12GB Rig
Citations and sources
- AMD — Ryzen AI product page
- TechPowerUp — GeForce RTX 3060 12 GB spec sheet
- Phoronix — 2026 GPU compute review
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.
