Qwen3.6 35B-A3B on a 12GB GPU with llama.cpp MTP: Direct Answer
As of September 2026, the "80 tok/s on a 12GB GPU" figure is real, but it was measured on an RTX 4070 Super, not an RTX 3060. In a llama.cpp MTP pull-request thread, the 4070 Super owner who produced it logged 95.6 tok/s with Qwen3.6-35B-A3B UD-Q4_K_XL and --spec-type draft-mtp. The same card ran at 79.8–97.0 tok/s on upstream llama.cpp and 110.24 tok/s on the ik_llama.cpp fork. An RTX 3060 12GB is much slower. One owner measured 20.5 tok/s without MTP and 31.5–34.5 tok/s with it, and a separate run without MTP measured 38.9 tok/s at UD-Q4_K_M. No 12GB card holds the model's weights: Unsloth's UD-Q4_K_M file is 22.13 GB, so every one of these setups keeps expert weights in system RAM with --n-cpu-moe.
Affiliate disclosure: SpecPicks earns commissions on qualifying Amazon purchases.
The full cross-card comparison: Best GPUs for Running Local LLMs in 2026 has the VRAM-tier table from 12 GB through 48 GB, with a source link on every row.
What changed — September 2026 refresh
- Corrected: the 80 tok/s headline. It came from an RTX 4070 Super, which logged 95.6 tok/s with MTP, not from an RTX 3060. Cited RTX 3060 runs land between 20.5 tok/s without MTP and 31.5–34.5 tok/s with it, and at 38.9 tok/s at UD-Q4_K_M. The title now states the cited figures.
- Corrected: MTP flags. The
--mtpand--mtp-draft-tokensflags never existed. The setup commands now use the merged--spec-type draft-mtpand--spec-draft-n-maxflags, per the llama-server README. - Corrected: file size and download. Unsloth's UD-Q4_K_M file is 22.13 GB, not ~19 GB, so no 12GB card holds it and
--n-cpu-moeexpert offload is now explicit. The non-publicQwen/Qwen3.6-35B-A3B-GGUFrepo was replaced with Unsloth's MTP GGUF. - Corrected: architecture. The model has 256 experts with 8 routed + 1 shared active per token, not "2 of 8", per the model card.
- Added: cited 12GB-class rows with active-parameter bandwidth ceilings for the RTX 3060, RTX 4070 Super, RTX 5070 and Arc B580. They include Gemma 4 26B-A4B at ~40–55 tok/s on an RTX 3060 and gpt-oss-20b at 155.87 tok/s on an RTX 5070.
- Removed: unsourced tables and claims. The first-person context-length table is gone, along with the uncited CPU-utilization, KV-cache and thermal claims. The context table now uses figures from the Qwen3.6-35B 12GB VRAM guide.
What's changed as of September 2026
This article first ran in May 2026, before llama.cpp's MTP support was merged. Six things have changed since then, and several of the original numbers were wrong:
- MTP is merged, and the flags are different. PR #22673 ("llama + spec: MTP Support") was merged on May 16, 2026. It is enabled with
--spec-type draft-mtpand--spec-draft-n-max N, per the llama-server README. The--mtpand--mtp-draft-tokensflags in the earlier version of this article do not exist. - The 80 tok/s headline came from a 4070 Super. The earlier version credited ~78 tok/s to a ZOTAC RTX 3060. No public benchmark supports that. The cited 3060 figures above top out near 39 tok/s.
- The model doesn't fit in 12 GB, and there is no official Qwen GGUF. The
Qwen/Qwen3.6-35B-A3B-GGUFrepository in the old download command is not a public Hugging Face repo. Community GGUFs are larger than the "~19 GB" claimed: UD-Q4_K_M is 22.13 GB, UD-Q4_K_S is 20.89 GB and UD-IQ4_XS is 17.73 GB. The MTP builds, which include the prediction head, add roughly 0.5 GB (22.66 GB at UD-Q4_K_M).--n-gpu-layers 999on its own does not work on a 12GB card. - MTP helps this MoE model less than it helps dense models. Unsloth's measurements show 1.4–2x for dense models but about 1.15–1.25x for MoE models. Unsloth recommends
--spec-draft-n-max 2, since acceptance falls from 83% to 50% at 4 draft tokens. On a 12GB card with heavy expert offload, MTP can be neutral or negative. One RTX 3060 owner running-ncmoe 99measured generation drop from ~26 to ~24 tok/s with MTP on. - The architecture description was wrong. Per the Qwen3.6-35B-A3B model card, the model has 40 layers, 256 experts and 8 routed + 1 shared expert active per token. It does not have "2 of 8" experts.
- Newer MoE options fit the same setup. Gemma 4 26B-A4B (25.2B total, 3.8B active) and gpt-oss-20b both fit 12GB cards at small quants. Rows are below.
Current 12GB-class rows (September 2026)
Each ceiling is the card's memory bandwidth divided by the bytes streamed per token. For MoE models that is approximated as the active-parameter share of the GGUF file (file size × active ÷ total parameters). Ceilings assume everything is resident in VRAM. With --n-cpu-moe, part of each token's expert reads come from system RAM instead, so the real limit is lower. Every measured figure below sits under its ceiling. Bandwidth figures follow from each card's bus width and memory clock on TechPowerUp: RTX 3060 12GB, 192-bit GDDR6 at 15 Gbps = 360 GB/s; RTX 4070 Super, 192-bit GDDR6X at 21 Gbps = 504 GB/s; RTX 5070, 192-bit GDDR7 at 28 Gbps = 672 GB/s; and Arc B580, 192-bit GDDR6 at 19 Gbps = 456 GB/s.
| GPU (12 GB) | Model (quant, HF file size) | Ceiling (active bytes) | Measured decode | Setup | Source |
|---|---|---|---|---|---|
| RTX 4070 Super | Qwen3.6-35B-A3B MTP UD-Q4_K_XL, 22.85 GB | ~257 tok/s (504 ÷ 1.96 GB) | 95.6 tok/s (single sample, 98% draft acceptance); 73.2 tok/s on a code benchmark prompt | upstream MTP branch, --spec-type draft-mtp --spec-draft-n-max 7, -fitt 1024, 131K ctx | PR #22673 comment |
| RTX 4070 Super | Qwen3.6-35B-A3B byteshape IQ4_XS, ~18.0 GB | ~326 tok/s (504 ÷ 1.54 GB) | 79.8–97.0 tok/s upstream MTP; 110.24 tok/s ik_llama.cpp | Ryzen 7 9700X, 48 GB DDR5-6000, 131K ctx, Q8 KV | Startup Fortune write-up of the r/LocalLLaMA post |
| RTX 4070 Super | Qwen3.6-35B-A3B APEX-I-Compact, ~16 GB | ~368 tok/s (504 ÷ 1.37 GB) | 64.0 tok/s at 32K ctx; 50.9 tok/s at 258K ctx | no MTP, --n-cpu-moe 16–22, i5-14600KF, 32 GB DDR4 | 12GB VRAM guide |
| RTX 3060 12GB | Qwen3.6-35B-A3B UD-Q4_K_M, 22.13 GB | ~190 tok/s (360 ÷ 1.90 GB) | 38.9 tok/s, flat through 8K ctx | no MTP, -ngl 99 -ncmoe 24 -fa 1 | InsiderLLM |
| RTX 3060 12GB | Qwen3.6-35B-A3B MTP Q8_0, 37.8 GB | ~111 tok/s (360 ÷ 3.24 GB) | 20.0–20.5 tok/s without MTP; 31.5–34.5 tok/s with MTP n-max 3 | -ngl 5 -ncmoe 32, q8_0 KV | PR #22673 comment |
| RTX 3060 12GB | Qwen3.6-35B-A3B UD-IQ3_XXS, 13.21 GB | ~318 tok/s (360 ÷ 1.13 GB) | 22.9 tok/s | no MTP, -ncmoe 25, 64K ctx, Xeon E5-2650 v4, 31 GB DDR4 | Jean Brito |
| RTX 3060 12GB | Gemma 4 26B-A4B UD-IQ2_M, 10.01 GB | ~238 tok/s (360 ÷ 1.51 GB) | ~40–55 tok/s | fully GPU-resident, 32K ctx, 3-bit "turbo3" KV cache (fork build) | julien9679 repo |
| RTX 5070 | gpt-oss-20b Q8_0, 12.11 GB | ~323 tok/s (672 ÷ 2.08 GB) | 155.87 tok/s (tg128) | llama.cpp b7083, Vulkan | Phoronix |
| Arc B580 | gpt-oss-20b Q8_0, 12.11 GB | ~219 tok/s (456 ÷ 2.08 GB) | 26.64 tok/s (tg128) | llama.cpp b7083, Vulkan (Mesa ANV) | Phoronix |
Active-parameter counts: Qwen3.6-35B-A3B is 35B total / 3B active, Gemma 4 26B-A4B is 25.2B / 3.8B, and gpt-oss-20b is 21B / 3.6B.
Gaps in the public data. No cited Qwen3.6-35B-A3B decode figure turned up for the plain RTX 4070 12GB, the RTX 5070 12GB, or the Arc B580. The one plain-4070 report in the MTP thread describes a regression on a development build, not a steady-state number. On the RTX 5070 and B580, gpt-oss-20b is the closest measured MoE proxy. The B580's 26.64 tok/s is far below its bandwidth ceiling on the Vulkan build Phoronix used.
What MTP is and why the gain is smaller on 12 GB
Multi-Token Prediction (MTP) is a form of speculative decoding. Qwen3.6 ships with an extra prediction head, trained with multi-step MTP per the model card, that drafts the next few tokens. The main model then verifies them in a single forward pass. llama.cpp loads that head from the same GGUF, so no separate draft model is needed (PR #22673). You need an MTP-enabled GGUF, such as Unsloth's MTP build. An ordinary GGUF has no head to draft with.
Two things shrink the gain on a 12GB card:
- MoE verification cost. Checking a batch of drafted tokens can route to different experts for each token. A sparse model therefore gains less from batching verification than a dense one does. Unsloth measured about 1.15–1.2x on average for the 35B-A3B versus 1.4x for dense 27B.
- VRAM spent on the head. The MTP head and its KV cache need VRAM that would otherwise hold expert layers. The PR author notes that prompt processing also slows down with MTP enabled. One 8GB owner found MTP forced fewer GPU layers and ended up at the same 32 tok/s as running without it.
The practical rule: turn MTP on, then compare --spec-draft-n-max values of 1, 2 and 3 against MTP off on your own card before settling on one.
Best Value: ZOTAC Gaming GeForce RTX 3060 Twin Edge OC 12GB (B08W8DGK3X)
The ZOTAC RTX 3060 Twin Edge OC 12GB is a dual-fan RTX 3060 with a 1777 MHz boost clock and 192-bit GDDR6 (360 GB/s). For Qwen3.6-35B-A3B it is a ~20–39 tok/s card, not an 80 tok/s one. The spread depends on quant, host CPU and whether MTP pays off, per the cited rows above. That is still comfortably interactive for chat. The price case has also changed: the catalog listing on its product page shows $499.99 at the time of this refresh, well above the $290–330 this article quoted in May.
Best alternative: MSI GeForce RTX 3060 Ventus 2X 12G (B08WRVQ4KR)
The MSI RTX 3060 Ventus 2X 12G uses the same GA106 chip and 12 GB of GDDR6, so the same tok/s rows apply. No public Qwen3.6 benchmark separates the two cards. Choose between them on price and cooler size. Its catalog listing on the product page also shows $499.99 at the time of this refresh.
If you are buying new for this workload, compare that price against a 12GB RTX 4070 Super. The 4070 Super is the card behind every 80+ tok/s figure above, with 504 GB/s of memory bandwidth against the 3060's 360 GB/s.
Companion CPU and system RAM: AMD Ryzen 7 5800X (B0815XFSGK)
With --n-cpu-moe, most expert weights sit in system RAM and the CPU computes those layers, so the host matters more here than in a fully GPU-resident setup. The cited 12GB setups ran on a Ryzen 7 9700X with 48 GB DDR5-6000 (the 80–110 tok/s 4070 Super runs), an i5-14600KF with 32 GB DDR4, and a 12-core Xeon E5-2650 v4 with 31 GB DDR4 (22.9 tok/s at IQ3_XXS). An 8-core AMD Ryzen 7 5800X with 32 GB of dual-channel DDR4 is a reasonable AM4 pairing. No public benchmark isolates its effect on this model, so treat CPU gains as varying by workload.
Size system RAM to the file. A 22 GB UD-Q4_K_M with roughly 10–11 GB on the GPU leaves about 11–12 GB of experts in RAM, plus the OS and runtime. That makes 32 GB the practical floor. Both cited 32 GB-class systems ran 64K contexts or more.
llama.cpp setup commands
Bring-up on a Linux host with a 12GB NVIDIA card and a current CUDA toolkit. Every flag below appears in the llama-server README:
Key flags explained:
--n-gpu-layers allwith--n-cpu-moe 24: put every layer on the GPU, then keep the MoE expert weights of the first 24 layers in system RAM. Raise N if you run out of memory and lower it if VRAM is left over. The 12GB guide above found the best N rises with context: 16 at 32K, 17 at 64K, 20 at 128K and 22 at 258K.-ot/--override-tensordoes the same job with a regex if you need finer control.- Alternatively, omit
--n-cpu-moeand let--fit(on by default) size the split automatically. The 95.6 tok/s 4070 Super run used-fitt 1024(a 1 GiB VRAM margin) instead of a manual N. --parallel 1: a single server slot. The default is auto, and extra slots multiply KV cache.--cache-type-k/v q8_0: halves KV-cache memory compared with the f16 default, leaving room for more experts on the GPU.--spec-type draft-mtp --spec-draft-n-max 2: enable MTP with two draft tokens. The llama-server default is 3; Unsloth recommends 2, and the 3060 owner cited above got the best results at 3. Measure with MTP off as well.
Throughput by context length (RTX 4070 Super 12GB)
Figures from the Qwen3.6-35B 12GB VRAM guide (RTX 4070 Super, i5-14600KF, 32 GB DDR4, ~16 GB APEX-I-Compact quant, no MTP):
| Context | KV cache | --n-cpu-moe | Decode tok/s | Prefill |
|---|---|---|---|---|
| 32K | q8_0 | 16 | 64.0 | not reported |
| 64K | q8_0 | 17 | 60.7 | 649 tok/s (45K-token input) |
| 128K | q8_0 | 20 | 55.3 | ~250 tok/s |
| 258K | q4_0 | 22 | 50.9 | 527 tok/s (45K-token input) |
The same guide reports that at 258K context, --n-cpu-moe values below 22 cause PCIe thrashing and drop decode to 20–25 tok/s. Offloading too few experts is worse than offloading too many.
When MTP doesn't help: edge cases
- Heavy expert offload. With nearly all experts in RAM (
-ncmoe 99), one 3060 owner measured ~26 tok/s without MTP and ~24 tok/s with it. - Too many draft tokens. Unsloth measured acceptance falling from 83% to 50% at 4 draft tokens. The 3060 owner's generation speed also dropped at n-max 4.
- Long prompts. The PR notes that prompt processing slows down with MTP enabled because of device-to-host embedding transfers (PR #22673).
- Vision input. A slot-corruption bug with MTP plus image input was reported on an RTX 5070 Ti (issue #22867, since closed). If you hit it on an older build, update llama.cpp or run with
--no-mmproj.
For other workloads the speedup varies by prompt. Compare MTP on and off on your own prompts.
Common pitfalls to avoid
- Assuming Q4 fits in 12 GB. It does not. UD-Q4_K_M is 22.13 GB and UD-Q5_K_M is 26.46 GB. Plan on
--n-cpu-moeor--fitat any 4-bit quant. Only quants of about 11–12 GB or less (UD-IQ2_M and below) come close to fitting entirely on the GPU. - Using a non-MTP GGUF with
--spec-type draft-mtp. The prediction head has to be in the file. Use an MTP build such as unsloth/Qwen3.6-35B-A3B-MTP-GGUF. - Using the old flag names.
--mtp,--mtp-draft-tokensand--draft-maxare not valid. The README marks--draft-maxas removed in favor of--spec-draft-n-max(llama-server README). - Leaving
--parallelon auto. Extra server slots each reserve KV cache and can push a 12GB card into out-of-memory errors at long contexts. - Too little system RAM. With 11+ GB of experts in RAM, 16 GB systems swap. The cited 12GB setups used 31–48 GB.
When NOT to buy a 3060 12GB for local LLM
If you want the 80–110 tok/s figures that circulate for this model, the RTX 3060 is the wrong card. Those came from an RTX 4070 Super with 504 GB/s of bandwidth, versus 360 GB/s on the 3060 (see the TechPowerUp specs linked above). Dense 70B-class models need far more memory than 12 GB at any usable quant. If you only run small dense models, a 12GB card is more than you need.
Citations and sources
- llama.cpp PR #22673, "llama + spec: MTP Support" (merged May 16, 2026): https://github.com/ggml-org/llama.cpp/pull/22673
- PR #22673 comment, RTX 4070 Super 12GB, 95.6 tok/s with draft-mtp: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4443537301
- PR #22673 comment, RTX 3060 12GB, 20.5 to 31.5–34.5 tok/s with MTP: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4393961295
- PR #22673 comment, RTX 3060 12GB with -ncmoe 99, MTP ~26 to ~24 tok/s: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4409603142
- PR #22673 comment, RTX 4070 12GB regression report: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4453841550
- PR #22673 comment, 8GB RTX 3060 Ti MTP VRAM cost: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4471492776
- PR #22673 comment, RTX 3060 laptop 6GB MTP draft sweep: https://github.com/ggml-org/llama.cpp/pull/22673#issuecomment-4375891776
- llama.cpp issue #22867, MTP with vision input: https://github.com/ggml-org/llama.cpp/issues/22867
- llama-server README (flag reference): https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- Startup Fortune, janvitos r/LocalLLaMA RTX 4070 Super benchmark (79.8–97.0 / 110.24 tok/s): https://startupfortune.com/110-toks-on-rtx-4070-super-with-qwen36-35b/
- InsiderLLM, Qwen3.6 35B on RTX 3060 12GB at 38.9 tok/s: https://insiderllm.com/guides/best-way-run-qwen-3-6-35b-moe-locally/
- Jean Brito, Qwen3.6 35B on RTX 3060 via -ncmoe (22.9 tok/s): https://jeanfbrito.github.io/posts/running-qwen3.6-35b-moe-rtx3060-ncmoe/
- Qwen3.6-35B 12GB VRAM guide (RTX 4070 Super context scaling): https://github.com/shiqikuangsan31/Qwen3.6-35B-12GB-VRAM-Guide
- Gemma 4 26B-A4B on RTX 3060 12GB: https://github.com/julien9679/Gemma-4-26B-A4B-on-RTX-3060-12GB
- Phoronix, llama.cpp Vulkan GPU comparison (gpt-oss-20b tg128): https://www.phoronix.com/review/llama-cpp-vulkan-eoy2025/4
- Unsloth Qwen3.6 guide (MTP speedups, draft-token advice): https://unsloth.ai/docs/models/qwen3.6
- Qwen3.6-35B-A3B model card: https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Unsloth Qwen3.6-35B-A3B GGUF file listing: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
- Unsloth Qwen3.6-35B-A3B MTP GGUF file listing: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF
- byteshape Qwen3.6-35B-A3B GGUF file listing: https://huggingface.co/byteshape/Qwen3.6-35B-A3B-GGUF
- am17an Qwen3.6-35BA3B MTP GGUF (37.8 GB Q8_0): https://huggingface.co/am17an/Qwen3.6-35BA3B-MTP-GGUF
- Gemma 4 26B-A4B model card: https://huggingface.co/google/gemma-4-26B-A4B-it
- Unsloth Gemma 4 26B-A4B GGUF file listing: https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF
- gpt-oss-20b model card: https://huggingface.co/openai/gpt-oss-20b
- Unsloth gpt-oss-20b GGUF file listing: https://huggingface.co/unsloth/gpt-oss-20b-GGUF
- TechPowerUp RTX 3060 12GB specs: https://www.techpowerup.com/gpu-specs/geforce-rtx-3060-12-gb.c3682
- TechPowerUp RTX 4070 Super specs: https://www.techpowerup.com/gpu-specs/geforce-rtx-4070-super.c4186
- TechPowerUp RTX 5070 specs: https://www.techpowerup.com/gpu-specs/geforce-rtx-5070.c4218
- TechPowerUp Arc B580 specs: https://www.techpowerup.com/gpu-specs/arc-b580.c4244
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.
Related guides
- Local LLM on RTX 3060 12GB: Why This Card Still Wins in 2026
- MTP Decoding on RTX 3060 12GB: When Multi-Token Prediction Helps
- AMD Ryzen AI Max+ 395 vs RTX 3060 12GB for Local LLM Inference
- Qwen3.6-35B-A3B vs Gemma4-26B-A4B: Which MoE Fits a 12GB RTX 3060
- Qwen 3.6 35B-A3B KV Cache Deep Dive
