Skip to main content

Qwen 3.6 35B-A3B on a 16GB GPU: RTX 5060 Ti vs RX 9060 XT Quant Table

A 16GB card holds the whole MoE at IQ3, while Q4 means expert offload. Here is the per-quant VRAM math and every public tok/s run side by side.

Qwen 3.6 35B-A3B fits a 16GB GPU at IQ3 (13.7 GB) and a public RTX 5060 Ti run hit 89 tok/s. Per-quant VRAM, offload math and an RX 9060 XT verdict.

Quick Answer

Can a 16GB GPU run Qwen 3.6 35B-A3B, and how fast? Yes. At UD-IQ3_S (13.7 GB) the whole model sits on the card, and the public RTX 5060 Ti 16GB run above holds 89 tok/s at 7.5K context and 46 tok/s at 108K (njannasch.dev). The RX 9060 XT 16GB has 320 GB/s of bandwidth to the 5060 Ti's 448 GB/s, and as of 2026-10-11 nobody has published a 35B-A3B run on it.

Qwen 3.6 35B-A3B on a 16GB GPU: RTX 5060 Ti vs RX 9060 XT Quant Table

Small-model (7–9B) reference throughput

These figures are for 7–9B models, not the 30–35B models this article covers. The table holds model size fixed so the cards can be compared with each other; for the size in the title, use the article's own figures.

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead. Rows marked for comparison are not the cards this article is about — they are included so the throughput column has something to be read against.

GPUVRAM Llama-3-8B class, Q4Street price Sources
NVIDIA GeForce RTX 5060 Ti 16 GB 71 tok/s11 runs · 8 sources $789street, all listings SpecPicks median of 11 runs; sources: LocalScore.ai (Mozilla Builders), RunAIHome, Runyard.dev, ComputingForGeeks +4 more
Radeon RX 9060 XT 16GB 16 GB 70.2 tok/s5 runs · 2 sources $530street, all listings SpecPicks median of 5 runs; sources: llama.cpp GitHub ROCm perf…, llama.cpp community benchmark…
NVIDIA GeForce RTX 3060for comparison 12 GB 55 tok/s23 runs · 9 sources $329MSRP SpecPicks median of 23 runs; sources: TYO Lab blog, Hardware Corner, llama.cpp GitHub (CUDA performanc…, Ajit Singh / Hardware-Corner +5 more

Which models fit on a RTX 5060 Ti?

Dense 30-35B models need about 19.9 GB for their Q4 weights; on the RTX 5060 Ti, the weights do not fit, so layers spill to system RAM and PCIe bandwidth sets the speed. The model this article covers is a mixture-of-experts design, so its file size at the quant the article tests, not this dense-class figure, decides whether it fits. RTX 5060 Ti carries 16 GB of VRAM. The weights column is the range of real Q4_K_M files on Hugging Face for the models in each size class (about 0.6 GB per billion parameters), and the runtime plus a usable context window wants about 2 GB on top; every tokens-per-second figure is a median over community-reported Q4 runs SpecPicks tracks for this card, with the run count and the source beside it.

Showing the model sizes this article covers and the band either side. Every size from 3B to 70B+, for every card SpecPicks tracks, is in the local-LLM GPU table.

Model size Weights at Q4 Fits in 16 GB? Measured Left for context Source
20-27B (Gemma 3 27B, Mistral Small)A 27B Q4_K_M file (about 16.7 GB) is more than a 16 GB card holds, so 16 GB means a Q3 quant or partial CPU offload; a 24B loads with a short context. 20 GB holds the class with a usable context window, 24 GB with a long one. about 16.7 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —
30-35B (Qwen 3 32B, QwQ 32B)The step change. A 24 GB card holds this entirely in VRAM; below that it is CPU offload. about 19.9 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —
70B+ (Llama 3.3 70B, Qwen 2.5 72B)One 48 GB card or two 24 GB cards. A 32 GB card runs it only with layers in system RAM. about 42.5 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —

Every RTX 5060 Ti benchmark run, with its source → GPU picks for local LLM — the same table across every card we track How we source these numbers

Yes. A 16GB card runs Qwen 3.6 35B-A3B entirely in VRAM at IQ3-class quants (13.2–13.7 GB files, per Unsloth's GGUF repo), and one public RTX 5060 Ti 16GB run logged 89 tok/s generation and 1,585 tok/s prefill at 7.5K context (njannasch.dev). Q4 needs expert offload. The RTX 5060 Ti's 448 GB/s and CUDA maturity give it the edge over the RX 9060 XT's 320 GB/s.

Introduction: why 16GB is the interesting tier for this MoE

Qwen 3.6 35B-A3B carries 35B total parameters but activates only about 3B per token: 8 routed experts plus 1 shared expert out of 256, across 40 layers, according to the official Qwen model card. That split makes the buying decision unusual. Total parameters set how much memory you need. Active parameters set how fast each token comes out.

At 8GB and 12GB, most of the model's experts have to live in system RAM, and the guides already on SpecPicks cover that offload-heavy path. At 24GB, a Q4 file fits with room to spare. 16GB is the one tier where a choice really changes the result: drop to an IQ3 quant and run the whole model on the GPU, or keep Q4 quality and send part of the experts to the CPU. The rest of this piece is about that choice, made on the two current-generation 16GB cards most buyers shortlist: the RTX 5060 Ti 16GB and the RX 9060 XT 16GB.

How the figures below were gathered. Every tok/s number comes from a public run: a blog write-up, a llama.cpp GitHub discussion, or a benchmark aggregator, linked on the row it supports. No first-party testing is reported here. Public runs use different llama.cpp builds, context lengths and KV-cache settings, so treat cross-row comparisons as directional. VRAM-at-context figures are computed from the model's published config and labeled as estimates.

Key Takeaways

  • The whole model fits at IQ3. UD-IQ3_XXS is 13.2 GB and UD-IQ3_S is 13.7 GB (Unsloth GGUF repo), small enough to run entirely on a 16GB card.
  • Full-GPU speed on the 5060 Ti is high. 89 tok/s generation and 1,585 tok/s prefill at 7.5K context with a q4_0 KV cache (njannasch.dev).
  • Q4 means offload. UD-Q4_K_M is 22.1 GB. A 16GB card with --cpu-moe expert offload lands around 25–35 tok/s in community reports (InsiderLLM).
  • The KV cache is unusually small. Only 10 of the 40 layers use full attention, with 2 KV heads (Qwen model card), so long context costs far less VRAM than it does on a dense 30B model.
  • There is no public RX 9060 XT run yet. On the closest MoE proxy, gpt-oss-20b, the 9060 XT posted 105.8 tok/s on Vulkan (llama.cpp discussion #15021) against 92.1 tok/s for the 5060 Ti (Hardware Corner).

How much VRAM does Qwen 3.6 35B-A3B need at each quant?

File sizes are taken from Unsloth's Qwen3.6-35B-A3B-GGUF repo. The "VRAM at 8K" column is an estimate: the file size plus about 0.5 GB for KV cache, compute buffer and DeltaNet state. The 0.5 GB allowance comes from the buffer breakdown in a measured RTX 5060 Ti run of the Qwen 3.5 version of this model (962 MiB of KV cache at 160K context, 258 MiB of compute buffer) (njannasch.dev), scaled down to 8K.

Quant (Unsloth)File sizeEst. VRAM at 8KFits fully on 16GB?Source
UD-Q2_K_XL12.3 GB~12.8 GBYesUnsloth
UD-IQ3_XXS13.2 GB~13.7 GBYesUnsloth
UD-IQ3_S13.7 GB~14.2 GBYes (the measured 5060 Ti config)Unsloth
UD-Q3_K_S15.4 GB~15.9 GBOnly on a headless card; no desktop headroomUnsloth
UD-Q3_K_M16.6 GB~17.1 GBNo, needs expert offloadUnsloth
UD-Q4_K_M22.1 GB~22.6 GBNo, about 7 GB of experts go to RAMUnsloth
UD-Q5_K_M26.5 GB~27 GBNo, heavy offloadUnsloth
UD-Q6_K29.3 GB~29.8 GBNo, heavy offloadUnsloth
Q8_036.9 GB~37.4 GBNo, needs 64GB system RAMUnsloth
BF1669.4 GB~70 GBNo, not practicalUnsloth

The practical cutoff is the Q3 line. Everything up to UD-IQ3_S leaves 1.5–2 GB free for long context and a desktop compositor. UD-Q3_K_S technically fits but leaves almost nothing. If you drive your monitors from the same card, take 0.5–1 GB off the budget before choosing.

How fast is it on the RTX 5060 Ti 16GB vs the RX 9060 XT 16GB?

Our catalog's baselines on a small dense model come first, because they show how close the two cards are before any MoE quirks apply. On 7–9B models at Q4, the RTX 5060 Ti 16GB has a median of 59.0 tok/s (n=11 runs, 7 sources, RTX 5060 Ti 16GB benchmarks), and the RX 9060 XT 16GB has a median of 61.1 tok/s (n=14 runs, 12 sources, RX 9060 XT 16GB benchmarks). On small dense models the two cards are effectively tied.

The public Qwen 3.6 35B-A3B data, plus the closest proxies:

CardModel / quantBackendPlacementPrefill tok/sGen tok/sSource
RTX 5060 Ti 16GBQwen 3.6 35B-A3B UD-IQ3_S, 7.5K ctxCUDA (llama.cpp b8838)Full GPU, q4_0 KV1,58589njannasch.dev
RTX 5060 Ti 16GBQwen 3.6 35B-A3B UD-IQ3_S, 108K ctxCUDAFull GPU, q4_0 KV—46njannasch.dev
RTX 5060 Ti 16GBQwen 3.5 35B-A3B UD-IQ3_XXS, 75K promptCUDAFull GPU, q4_0 KV1,12347–51njannasch.dev
RTX 5060 Ti / 4060 Ti 16GBQwen 3.6 35B-A3B IQ3_XS, 8–16K ctxCUDAFull GPU—~32–40 (community)InsiderLLM
RTX 5070 Ti 16GBQwen 3.6 35B-A3B UD-Q4_K_MCUDA--cpu-moe offload, 32GB RAM—~25–35 (community)InsiderLLM
RX 9060 XT 16GBQwen 3.6 35B-A3B——No public run foundNo public run found—
RX 9060 XT 16GBgpt-oss-20b MXFP4 (MoE proxy)VulkanFull GPU3,464105.8llama.cpp #15021
RTX 5060 Ti 16GBgpt-oss-20b MXFP4 (MoE proxy), 4K ctxCUDAFull GPU3,58592.1Hardware Corner

Read the table this way:

  • The 89 tok/s figure is real but best-case. It is one owner's run at short context with an aggressive q4_0 KV cache. The same author saw a 14× prefill collapse (113 vs 1,585 tok/s) when switching the KV cache to q5_1/q4_1 (njannasch.dev), so KV settings matter as much as the card does.
  • The community 32–40 tok/s range is the safer planning number for a stock setup with less tuning.
  • The RX 9060 XT has no direct data. On gpt-oss-20b, a different MoE with about 3.6B active parameters, it matches or beats the 5060 Ti on generation. The two rows come from different runs and settings, though, and gpt-oss fits with far more room to spare. Expect the 9060 XT to land near the 5060 Ti at full-GPU IQ3. Bandwidth predicts it will fall somewhat behind once context grows and every token reads more KV.

Spec delta: what separates the two 16GB cards for MoE inference?

SpecRTX 5060 Ti 16GBRX 9060 XT 16GBRTX 3060 12GB (reference)Why it matters
VRAM16 GB GDDR716 GB GDDR612 GB GDDR6Sets which quant fits without offload
Memory bus128-bit128-bit192-bitWidth × data rate = bandwidth
Bandwidth448 GB/s320 GB/s360 GB/sThe main limit on generation tok/s
Board power180 W160 W170 WPSU sizing; all three run on 550–650W units
PCIe5.0 x85.0 x164.0 x16Matters for offloaded prefill on older boards
Launch MSRP$429$349$329Perf-per-dollar at launch
LLM backendCUDA (mature)ROCm 7.x / VulkanCUDADay-one support for new llama.cpp kernels

Spec sources: TechPowerUp RTX 5060 Ti 16GB, TechPowerUp RX 9060 XT 16GB, TechPowerUp RTX 3060 12GB.

The 5060 Ti has 40% more bandwidth than the 9060 XT (448 vs 320 GB/s). On a model whose generation speed depends on how fast it streams active weights and KV, that gap is the deciding spec. The 9060 XT's advantages are its $80-lower launch MSRP and a full x16 link. The x16 link matters if you put the card in a PCIe 4.0 board and offload experts: the 5060 Ti's x8 link drops to PCIe 4.0 x8 there.

What happens when the model spills past 16GB?

For Q3_K_M and anything larger, llama.cpp's --n-cpu-moe N keeps the expert tensors of the first N layers in system RAM, and --cpu-moe keeps all of them there. Attention, the shared expert, embeddings and KV cache stay on the GPU. Both flags are in current llama.cpp builds.

What owners report:

  • Generation drops less than you'd expect. The CPU only reads the 8 routed experts each token selects, so a 16GB card on UD-Q4_K_M with full --cpu-moe still lands around 25–35 tok/s (InsiderLLM). Raise N only as far as you need to stop the out-of-memory errors. Each layer you keep on the GPU wins back some speed.
  • Prefill drops more. Long prompts push large batches through the offloaded experts, and that becomes CPU- and PCIe-bound. Expect prompt processing to fall well below the 1,585 tok/s full-GPU figure.
  • Partial layer offload is the worst option. One RTX 3080 16GB owner who split whole layers with --n-gpu-layers 24 on Vulkan saw 14–22 tok/s on UD-Q4_K_XL. That rose to about 30 tok/s once they switched to the MTP build (The Autodidacts). Use expert offload, not layer offload.

System RAM: 32GB is the floor for a Q4 or Q5 file with experts offloaded (InsiderLLM). 64GB covers Q6/Q8 experiments with a browser and IDE open. Run dual-channel memory: offloaded experts depend on RAM bandwidth, and a single-DIMM budget prebuilt roughly halves it.

How does context length change the picture?

This is where Qwen 3.6's hybrid design helps. Thirty of the 40 layers are Gated DeltaNet linear-attention layers with a fixed-size state. Only 10 are full-attention layers with a growing KV cache, and those use 2 KV heads at a head dimension of 256 (Qwen model card). Working from that config, the FP16 KV cache costs about 20 KB per token: 10 layers × 2 (K and V) × 2 heads × 256 × 2 bytes. That puts it at:

ContextFP16 KV (est.)q8_0 KV (est.)q4_0 KV (est.)Largest fully-on-GPU quant (est.)
8K~0.16 GB~0.09 GB~0.05 GBUD-Q3_K_S (headless) / UD-IQ3_S
32K~0.66 GB~0.35 GB~0.18 GBUD-IQ3_S
64K~1.3 GB~0.7 GB~0.37 GBUD-IQ3_S (tight at FP16 KV)
262K (native max)~5.4 GB~2.8 GB~1.5 GBUD-IQ3_S with q4_0 KV only

The measured run backs the arithmetic up. The RTX 5060 Ti kept UD-IQ3_S plus the full 262K native context on the card with a q4_0 KV cache, with 756 MB of VRAM left free (njannasch.dev). Context, in other words, barely limits you on 16GB. Weight size does. For the quality side of the q8_0 vs q4_0 KV choice, see our Qwen 3.6 35B-A3B KV-cache quantization guide. Keep in mind that the measured q4_0 prefill speed beat the q5_1/q4_1 mix by 14× on this card.

Is 16GB worth it over the RTX 3060 12GB for this model?

The 12GB path is well documented. With --n-cpu-moe expert offload, an RTX 3060 12GB holds about 38 tok/s (InsiderLLM), and the MTP build pushes it higher (our 12GB MTP guide). The budget cards for that route are the MSI Gaming GeForce RTX 3060 12GB and the ZOTAC RTX 3060 Twin Edge OC 12GB.

What 16GB buys you:

  1. Running the whole model on the GPU. No 12GB card can hold even UD-Q2_K_XL (12.3 GB) plus buffers, so the 3060 always offloads. On 16GB, IQ3 runs entirely on the card, with full-GPU prefill to match.
  2. Prefill. 1,585 tok/s full-GPU prefill on the 5060 Ti (njannasch.dev) means a 30K-token document is ingested in about 20 seconds. Offloaded prefill on a 12GB card takes noticeably longer.
  3. Long context at low cost. 262K native context fits on the card.

Perf-per-dollar at launch MSRP, using the planning figures above: the RTX 3060 12GB gives about 38 tok/s for $329 (about 0.12 tok/s per dollar). The RTX 5060 Ti 16GB gives 32–40 tok/s in community reports and up to 89 tok/s tuned, for $429 (0.07–0.21 tok/s per dollar). The RX 9060 XT at $349 can't be scored until someone publishes a 35B-A3B run.

Which 16GB card should you buy for Qwen 3.6 35B-A3B?

Get the RTX 5060 Ti 16GB if… you want the configuration that has actually been measured. It is the only 16GB card with a public 35B-A3B run (89 tok/s full-GPU at IQ3). It has the most bandwidth in this group at 448 GB/s, and new llama.cpp features like MTP and new quant kernels reach CUDA first. The ASUS Dual GeForce RTX 5060 Ti 16GB OC is the catalog pick.

Get the RX 9060 XT 16GB if… you're on Linux, comfortable with ROCm 7.x or Vulkan, and price matters more than the last 20–30% of generation speed. Its launch MSRP is $80 lower, its PCIe 5.0 x16 link helps offloaded prefill on older boards, and its gpt-oss-20b MoE result shows the card handles sparse models well. Consider the ASRock Radeon RX 9060 XT Challenger 16GB OC or the ASUS Dual Radeon RX 9060 XT 16GB.

Stay on an RTX 3060 12GB if… you already own one, your prompts are under about 8K tokens, and roughly 38 tok/s with expert offload is enough. A 16GB card gives you full-GPU prefill and headroom. It does not give you a better model.

Recommended pick: the RTX 5060 Ti 16GB running UD-IQ3_S with a q4_0 KV cache and full GPU placement. It is the one 16GB setup with published numbers for this exact model, and those numbers are strong at short context while holding up at 100K+. If you need Q4 quality, keep the same card, add 32GB of dual-channel RAM, and use --n-cpu-moe.

Common pitfalls

  • Using layer offload instead of expert offload. --n-gpu-layers below the maximum puts entire layers on the CPU, attention included. Keep -ngl 99 and offload experts with --n-cpu-moe instead.
  • Mixing KV cache types. The 14× prefill penalty for q5_1/q4_1 versus q4_0 on the 5060 Ti (njannasch.dev) is the biggest single gotcha. Benchmark your KV setting before blaming the card.
  • Leaving no room for the desktop. UD-Q3_K_S fits on paper but not on a card that also drives two 1440p monitors.
  • Single-channel RAM with offload. Offloaded experts depend on RAM bandwidth. One DIMM makes offloading close to worthless.
  • Trusting calculator estimates. Some "can I run it" sites list Q4_K_M on the 9060 XT at a few tok/s from formulas, not measurements. Check whether a number was measured before you rely on it.

When NOT to buy a 16GB card for this model

If you need Q5 or higher quality without offload, 16GB is the wrong tier. Q5_K_M is 26.5 GB (Unsloth), so look at 24–32GB cards instead (our 35B-A3B vs 27B dense 24GB comparison). If you serve several concurrent users, expert offload thrashes. And if you already own a 12GB card and work at short context, the upgrade buys less than its price suggests.

Bottom line

A 16GB card is the cheapest tier where Qwen 3.6 35B-A3B runs entirely on the GPU, but only at IQ3-class quants. The RTX 5060 Ti 16GB is the measured, lower-risk choice thanks to its bandwidth lead and mature CUDA stack. The RX 9060 XT 16GB is the value pick for Linux users willing to wait for public numbers on this model. If you already own an RTX 3060 12GB and work at short context, it remains a good host for this model.

Live price comparison

For current prices on the two main cards, open the ASUS Dual RTX 5060 Ti 16GB vs ASRock RX 9060 XT Challenger 16GB comparison. Its live price buttons show today's Amazon prices. Prices change often, so check them before you buy.

Citations and sources

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Amazon & eBay listings, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

I'm still mad… but buy it anyway - RTX 3060 Review — Linus Tech Tips on YouTube

Frequently asked questions

Does Qwen 3.6 35B-A3B fit entirely in 16GB of VRAM?
Yes at IQ3-class quants. Unsloth's UD-IQ3_XXS (13.2 GB) and UD-IQ3_S (13.7 GB) files fit on a 16GB card with room for KV cache, and a public RTX 5060 Ti run kept UD-IQ3_S plus the full 262K native context on-GPU with a q4_0 KV cache. UD-Q3_K_M (16.6 GB) and UD-Q4_K_M (22.1 GB) do not fit and need expert offload to system RAM.
Why does an MoE model run reasonably even when part of it sits in system RAM?
Only about 3B parameters are active per token, so each generated token reads a small slice of the weights. When llama.cpp offloads whole expert tensors to the CPU, the attention layers and shared weights stay on the GPU and generation slows less than it would for a dense 35B model. Prefill still drops noticeably, so long prompts feel the offload penalty more than short chat turns do.
Is ROCm or Vulkan better for the RX 9060 XT on this model?
Both backends run Qwen 3.6 MoE GGUFs in recent llama.cpp builds. Public results vary by driver version and by whether flash attention is enabled, so the article reports each backend in its own row with its source instead of declaring a single winner. Linux users on a current ROCm release tend to report the most consistent prefill numbers, while Vulkan is the simpler path on Windows.
How much system RAM do I need if I offload experts?
Plan for at least 32GB of system RAM when offloading part of a Q4 or Q5 file. 64GB leaves room for Q6/Q8 experiments and a browser running alongside. Dual-channel DDR5 helps, because offloaded experts are bound by memory bandwidth. Single-channel configurations, which are common in budget prebuilts, can cut offloaded generation speed sharply.
Should I just keep my RTX 3060 12GB for this model?
If you already own an RTX 3060 12GB, the published 12GB guides show that the model is usable with expert offload at Q4. A 16GB card mainly buys longer context and less offloading, not a different class of model. Upgrade if you regularly need 32K or more of context or faster prefill on long documents. Otherwise the 3060 remains a reasonable host for this MoE.

— Mike Perry · Updated 2026-10-11

Parts this article names

Amazon Associate — prices tracked 2026-10-11, may vary.