Gemma 4 31B on a 16GB GPU: RTX 5060 Ti vs RX 9060 XT vs Arc A770 Quant Table (2026)
Q3_K_M is the largest Gemma 4 31B quant that stays on a 16GB card. Here is how the RTX 5060 Ti, RX 9060 XT and Arc A770 compare on it.
By Mike Perry, Founder & Editor-in-Chief · Published 2026-10-10 · Updated 2026-10-10 · 17 min read
Can a 16GB GPU run Gemma 4 31B? Per-quant VRAM table, KV-cache math and sourced tok/s for the RTX 5060 Ti, RX 9060 XT and Arc A770 vs a 12GB RTX 3060.
Quick Answer
Yes, a 16GB GPU can run Gemma 4 31B fully on the card, but only at 3-bit quants. Unsloth's Q3_K_M file is 14.74 GB, and on an RTX 5060 Ti 16GB it peaked at 15.0 GB of VRAM with an 8K q4_0 KV cache, per njannasch.dev's Gemma 4 on a 5060 Ti test. Q4_K_M (18.32 GB on the unsloth GGUF listing) does not fit. Pick the ASUS Dual RTX 5060 Ti 16GB. Its 448 GB/s of bandwidth and the CUDA backend give it the only published 16GB run, at 26–29 tok/s on UD-IQ3_XXS.
This page is for owners of 16GB cards, and buyers considering one, who have already tried Gemma 4 31B on 12GB and hit the offload wall. The 12GB RTX 3060 Gemma 4 31B guide covers the 2-bit and heavy-offload end of the range. This page covers the next tier up. Three 16GB cards are worth considering in 2026: NVIDIA's RTX 5060 Ti 16GB, AMD's RX 9060 XT 16GB and Intel's Arc A770 16GB. The RTX 3060 12GB is included as the baseline most readers are upgrading from. The format follows the Qwen 3.6 27B 16GB quant table, but Gemma 4 31B is a larger, differently built model, so the conclusions differ. Every VRAM figure below is either a published measurement or calculated from the model's own config, and each one says which.
Small-model (7–9B) reference throughput
These figures are for 7–9B models, not the 30–35B models this article covers. The table holds model size fixed so the cards can be compared with each other; for the size in the title, use the article's own figures.
Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead.
The 30-35B class this article is about needs about 19.9 GB for its Q4 weights; on the RTX 5060 Ti, the weights do not fit, so layers spill to system RAM and PCIe bandwidth sets the speed. RTX 5060 Ti carries 16 GB of VRAM. The weights column is the range of real Q4_K_M
files on Hugging Face for the models in each size class (about 0.6 GB per billion parameters),
and the runtime plus a usable context window wants about 2 GB on top; every tokens-per-second figure is a
median over community-reported Q4 runs SpecPicks tracks for this card, with the run count and
the source beside it.
Showing the model sizes this article covers and the band either side. Every size from 3B to 70B+, for every card SpecPicks tracks, is in the local-LLM GPU table.
Model size
Weights at Q4
Fits in 16 GB?
Measured
Left for context
Source
20-27B (Gemma 3 27B, Mistral Small)A 27B Q4_K_M file (about 16.7 GB) is more than a 16 GB card holds, so 16 GB means a Q3 quant or partial CPU offload; a 24B loads with a short context. 20 GB holds the class with a usable context window, 24 GB with a long one.
about 16.7 GB
Nospills to system RAM — PCIe bandwidth sets the speed
—
none
—
30-35B (Qwen 3 32B, QwQ 32B)The step change. A 24 GB card holds this entirely in VRAM; below that it is CPU offload.
about 19.9 GB
Nospills to system RAM — PCIe bandwidth sets the speed
—
none
—
70B+ (Llama 3.3 70B, Qwen 2.5 72B)One 48 GB card or two 24 GB cards. A 32 GB card runs it only with layers in system RAM.
about 42.5 GB
Nospills to system RAM — PCIe bandwidth sets the speed
The largest quant that runs fully on 16GB is Q3_K_M (14.74 GB), and only with a quantized KV cache. The published 5060 Ti run peaked at 15.0 GB at 8K context with q4_0 KV (njannasch.dev).
UD-IQ3_XXS (11.84 GB) is the comfortable 16GB quant. It generated 26–29 tok/s on an RTX 5060 Ti and held 65K context with q4_0 KV, per the same test.
Offload is a cliff, not a slope. A Q6_K file (23.4 GB) on the same card measured 0.62 tok/s generation in llama.cpp issue #26674.
Gemma 4's sliding-window layers keep long context cheap. 50 of the 60 layers attend to only 1,024 tokens, per the Gemma 4 31B config. Most of the KV cache stays a fixed size no matter how long the context gets.
Bandwidth sets the speed ceiling: 560 GB/s on the Arc A770, 448 GB/s on the RTX 5060 Ti and 320 GB/s on the RX 9060 XT (TechPowerUp RTX 5060 Ti, TechPowerUp RX 9060 XT, TechPowerUp Arc A770). Software maturity decides how much of that ceiling each card reaches.
The VRAM column is calculated. It adds the file size in GiB, the 8K f16 KV cache derived from the model config (1.8 GiB, explained in the KV section below), and about 0.5 GiB of llama.cpp compute buffer. The fit test is against roughly 15.5 GiB, the usable memory per RTX 5060 Ti 16GB after the driver reserve, as reported in LLMKube's dual-5060 Ti bake-off. This method lands within 0.1–0.3 GB of the two measured peaks in the njannasch.dev test. "Layers offloaded" is the excess divided by an average per-layer size (file ÷ 60 layers), rounded up.
Quant
GGUF file size
Est. VRAM at 8K (f16 KV)
Fits fully in 16GB at 8K?
Layers offloaded (of 60)
Quality note
UD-IQ3_XXS
11.84 GB
13.3 GiB (measured peak 13.2 GB)
Yes, room for 32K–65K with q4_0 KV
0
Lowest 3-bit; the speed pick
Q2_K
12.63 GB
14.1 GiB
Yes
0
Larger than IQ3_XXS but lower quality; skip it
Q3_K_S
13.21 GB
14.6 GiB
Yes
0
Safe 3-bit with f16 KV at 8K
Q3_K_M
14.74 GB
16.0 GiB (measured 15.0 GB with q4_0 KV)
Partial with f16 KV; yes with q8_0/q4_0 KV
0 with q4_0 KV, ~3 with f16 KV
Best quality that stays on-card
IQ4_XS
16.37 GB
17.5 GiB
No
~9
First 4-bit option
Q4_K_M
18.32 GB
19.4 GiB
No
~14
The usual default; needs 24GB
Q5_K_M
21.66 GB
22.5 GiB
No
~21
24GB tier
Q6_K
25.20 GB
25.8 GiB
No (measured 0.62 tok/s offloaded)
~27
32GB tier
Q8_0
32.64 GB
32.7 GiB
No
~34
48GB tier
BF16
61.41 GB
59.5 GiB
No
~47
Reference weights
Sources: file sizes from the unsloth, bartowski and ggml-org GGUF listings above. Measured peaks are from njannasch.dev. The 0.62 tok/s Q6_K figure is from llama.cpp issue #26674. Google also ships an official QAT build. The 17.65 GB Q4_0 file in google/gemma-4-31B-it-qat-q4_0-gguf is trained to hold quality at 4-bit, but it is still about 2 GB too large for a 16GB card.
Two caveats. First, loading the vision projector (mmproj, 1.2 GB) for image input adds about 1.1 GiB to every row, and that alone drops Q3_K_M from "fits" to "offload." Second, if the same card drives your desktop, subtract a few hundred MB from the 15.5 GiB budget.
Step 0: is 16GB actually your bottleneck?
Before you buy a card, check which constraint you are actually hitting. Answer three questions.
1. Are you offloading now, and by how much? If your current setup runs Gemma 4 31B Q4_K_M with 10+ layers in system RAM, a 16GB card does not fix that. Q4_K_M needs about 19.4 GiB at 8K, so it will still offload around 14 layers. The upgrade only helps if you are willing to drop to a 3-bit quant that fits completely.
2. Is Q3 on-card better than Q4 with offload? For interactive use, yes, by a wide margin. Every offloaded layer is read from system RAM over PCIe on every token. The issue #26674 report shows how bad that gets: 23.4 GB of Q6_K on a 16GB 5060 Ti generated 0.62 tok/s while prompt processing ran at 41.84 tok/s. The fully resident Q3_K_M ran at 17–19 tok/s on the same card model. Q4 with offload is reasonable only for batch jobs nobody is waiting on.
3. Is a 12GB RTX 3060 with offload good enough? On 12GB, the largest on-card option is about UD-IQ2_M (10.75 GB) with q4_0 KV. The calculated total of about 11.0 GiB leaves almost no room for context. If 31B is something you run occasionally and you can tolerate 2-bit quality, keep the 3060. If 31B is your daily model, the 3-bit tier is a real quality step and needs 16GB.
Worked example A: daily coding chat at 8K. Q3_K_M with --cache-type-k q4_0 --cache-type-v q4_0 and flash attention fits a headless 16GB card. Expect the 17–19 tok/s the njannasch.dev run measured on a 5060 Ti.
Worked example B: long-document Q&A at 64K. Drop to UD-IQ3_XXS with q4_0 KV. The same source held 65K context at a 13.8 GB peak and still generated 18 tok/s at 50K tokens.
Worked example C: you need 4-bit quality. No 16GB card does this without offload. Look at 24GB instead, or use Gemma 4's 26B MoE sibling, which the njannasch.dev test ran 3.5x faster on the same card.
How much VRAM does the KV cache add at 8K, 32K and 128K context?
Gemma 4 31B's config.json defines 60 layers. 50 are sliding-window layers (16 KV heads, head size 256, 1,024-token window) and 10 are full-attention layers (4 KV heads, head size 512). The two types use memory very differently:
Sliding layers: 800 KiB per token in f16, but llama.cpp sizes this cache to the window plus one micro-batch (1,024 + 512 = 1,536 cells by default). That is a fixed ~1.17 GiB in f16, no matter how long the context is.
Full-attention layers:80 KiB per token in f16, growing linearly with context.
Context
Fixed sliding cache (f16 / q8_0 / q4_0)
Growing global cache (f16 / q8_0 / q4_0)
Total KV (f16 / q8_0 / q4_0)
8K
1.17 / 0.62 / 0.33 GiB
0.62 / 0.33 / 0.18 GiB
1.80 / 0.95 / 0.51 GiB
32K
1.17 / 0.62 / 0.33 GiB
2.50 / 1.33 / 0.70 GiB
3.67 / 1.95 / 1.03 GiB
128K
1.17 / 0.62 / 0.33 GiB
10.0 / 5.31 / 2.81 GiB
11.2 / 5.94 / 3.14 GiB
These figures are calculated from config.json, assuming llama.cpp stores K and V separately for every layer. The config's attention_k_eq_v flag could let a runtime halve the global portion, but the measured peaks above are consistent with the conservative figure.
What this means on 16GB:
Do not pass --swa-full. That flag makes the sliding layers cache the full context too. At 32K that is 25 GiB of f16 cache, and it is the most common reason Gemma 4 runs out of memory on 16GB cards, per njannasch.dev.
q8_0 KV is the default to reach for. It roughly halves the cache, which is the margin that lets Q3_K_M stay on-card. q4_0 saves more and is what the 65K run used, but it costs some long-context recall quality.
128K is out of reach for the dense 31B on 16GB. Even IQ3_XXS with q4_0 KV comes to about 14.7 GiB in this model. The njannasch.dev 131K configuration loaded at a 15.5 GB peak but ran out of memory before reaching 108K tokens of generation.
Spec-delta table: RTX 5060 Ti 16GB vs RX 9060 XT 16GB vs Arc A770 16GB vs RTX 3060 12GB
Card
VRAM
Memory bandwidth
Launch MSRP
Board power
llama.cpp backend
GeForce RTX 5060 Ti 16GB
16GB GDDR7, 128-bit
448 GB/s
$429
180W
CUDA
Radeon RX 9060 XT 16GB
16GB GDDR6, 128-bit
320 GB/s
$349
160W
ROCm (HIP) or Vulkan
Arc A770 16GB
16GB GDDR6, 256-bit
560 GB/s
$349
225W
SYCL or Vulkan
GeForce RTX 3060 12GB
12GB GDDR6, 192-bit
360 GB/s
$329
170W
CUDA
Specs and launch prices are from TechPowerUp's database entries: RTX 5060 Ti 16 GB, RX 9060 XT 16 GB, Arc A770 and RTX 3060 12 GB. These are launch prices, all at MSRP. Street prices move weekly; the live price buttons on this page carry today's numbers.
All three 16GB cards hold exactly the same quants, so the per-quant table applies to each. They differ in how fast they read those weights. On paper, the Sparkle Arc A770 ROC OC 16GB has the widest bus of the four, 256-bit at 560 GB/s, which is 25% more than the 5060 Ti and 75% more than the 9060 XT. Whether you get that bandwidth in practice depends on the backend, as the next section shows.
Which 16GB card generates Gemma 4 31B tokens fastest?
As of October 2026, only one card has a published single-card Gemma 4 31B run that fits on the card: the RTX 5060 Ti. For the others, the table gives the bandwidth ceiling. The ceiling is memory bandwidth divided by file size, a physical upper bound rather than a measurement.
The 5060 Ti reached about 77% of its ceiling on IQ3_XXS and about 62% on Q3_K_M with q4_0 KV, where dequantizing the cache adds work. If the RX 9060 XT reached the same 77%, it would land around 21 tok/s on IQ3_XXS. Treat that as a projection until someone publishes the run.
To compare how much of their bandwidth each card actually delivers in real software, here is the median of each card's published 7–9B generation runs (llama.cpp and Ollama) from the SpecPicks benchmark pages:
Read these medians with care. Two of the three RX 9060 XT runs use Llama 2 7B q4_0, a smaller file than the Llama 3.1 8B Q4_K_M behind most of the NVIDIA runs, so the 9060 XT median is flattered. The A770's published runs span 15–67 tok/s depending on backend and build. In the llama.cpp Vulkan scoreboard (#10879), the A770 posts 52.6 tok/s on Llama 2 7B q4_0, against 70.5 for the RX 9060 XT and 75.9 for the RTX 3060 on the same harness. The card with the most bandwidth on paper finishes last on the same benchmark. Scaling to 31B does not change that pattern, because generation stays bandwidth-bound and backend efficiency carries over.
Prefill vs generation: why prompt processing favours the RTX 5060 Ti's CUDA path
Generation reads every weight once per token, so it is limited by bandwidth. Prefill processes the whole prompt in parallel, so it is limited by compute and kernel quality. It decides how long you wait for the first token, and with a 31B model and long pastes, that wait is the delay you notice.
On the same LocalScore harness, the 5060 Ti prefills 58% faster than the 3060. On the Vulkan scoreboard, the 9060 XT prefills 18% faster than the 3060, and the A770 runs at about half the 9060 XT's rate. Chaining the two ratios through the RTX 3060 puts the 5060 Ti ahead of the 9060 XT on prefill. That is indirect, since the harnesses differ, but it lines up with Blackwell's 5th-generation tensor cores and the maturity of the CUDA flash-attention kernels.
Gemma 4's architecture also helps prefill. The 50 sliding layers attend to only 1,024 tokens, so attention cost on those layers stops growing with prompt length. The 10 global layers carry the long-context cost.
Backend gotchas per card
RTX 5060 Ti 16GB (CUDA). Use a llama.cpp build compiled for compute capability 12.0 (Blackwell). Older prebuilt binaries fall back to PTX JIT and start slowly. The most-missed step: check that -ngl 99 actually put every layer on the GPU. In issue #26674, the reporter saw 100% CPU and under 15% GPU utilization during Gemma 4 generation. A maintainer attributed the slow speed to an oversized dense model spilling into system RAM. If nvidia-smi shows less VRAM in use than your file size, layers are on the CPU.
RX 9060 XT 16GB (ROCm or Vulkan). Vulkan is the easier install and works on Windows. ROCm on Linux needs a release that lists RDNA 4 (gfx1200) as supported. The most-missed step: on ROCm, set HSA_OVERRIDE_GFX_VERSION only if your ROCm build predates official gfx1200 support. On a current build it does more harm than good. The parsapp RX 9060 XT benchmark repo found Vulkan slightly ahead on generation at 14B, so start there.
Arc A770 16GB (SYCL or Vulkan). Intel's IPEX-LLM repository was archived in January 2026, so the guides that start with "install IPEX-LLM" are out of date. Use llama.cpp's own Vulkan or SYCL backend. The most-missed step: try Vulkan first. In llama.cpp issue #19918, an A770 owner measured SYCL at 10 tok/s generation versus 68 on Vulkan with the same MoE models. Also enable Resizable BAR in the BIOS. Arc cards lose significant performance without it.
Common pitfalls on all three cards:
Leaving the vision projector loaded.--mmproj adds about 1.2 GB. Load it only when you send images.
Using --swa-full from an old guide. It undoes Gemma 4's cache savings (see the KV section).
Running an outdated build. Gemma 4 support, including its chat template and iSWA cache, needs a 2026 llama.cpp build.
Benchmarking with the browser open. Hardware-accelerated browsers can hold several hundred MB of VRAM, and on 16GB that margin is what Q3_K_M needs.
Perf-per-dollar and perf-per-watt math (at MSRP)
All prices are launch MSRP from the spec table, so there is one consistent price basis. Power is rated board power, not measured draw. Two metrics are shown: the measured 7–9B median, and the calculated Gemma 4 31B IQ3_XXS bandwidth ceiling.
Card
7–9B median tok/s per $100 MSRP
7–9B median tok/s per 100W
31B IQ3_XXS ceiling per $100 MSRP
31B IQ3_XXS ceiling per 100W
RTX 5060 Ti 16GB
13.8
32.9
8.8
21.0
RX 9060 XT 16GB
19.4
42.2
7.7
16.9
Arc A770 16GB
10.6
16.4
13.6
21.0
RTX 3060 12GB
17.4
33.7
n/a (does not fit)
n/a
The two metrics point in different directions. On measured small-model throughput, the RX 9060 XT gives the most tokens per launch dollar and per watt, although its median is flattered by smaller test files. On the bandwidth ceiling for this specific 31B model, the A770 leads per dollar, and the A770 and 5060 Ti tie per watt. In practice, the 5060 Ti is the only card with a measured 31B result (26–29 tok/s), and that measurement counts for more than either ceiling.
Verdict matrix
Get the RTX 5060 Ti 16GB if…
Gemma 4 31B is your daily model. It is the only 16GB card with a published on-card run (26–29 tok/s on IQ3_XXS, 17–19 tok/s on Q3_K_M).
You paste long prompts and want the fastest prefill of the group.
You want CUDA's ecosystem: llama.cpp, vLLM, ExLlama and Ollama all support Blackwell.
Get the RX 9060 XT 16GB if…
Most of your work is 7–14B models, where its measured throughput matches or beats the 5060 Ti for less money at launch.
You accept roughly 29% less bandwidth on 31B (320 vs 448 GB/s) in exchange for 160W board power.
Buy the ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC and run Gemma 4 31B Q3_K_M with q8_0 or q4_0 KV cache at 8K. Switch to UD-IQ3_XXS with q4_0 KV when you need 32K–65K context. The deciding spec is memory bandwidth backed by mature software. The 5060 Ti's 448 GB/s delivered a measured 77% of its ceiling on this exact model, while the A770's larger paper ceiling has no 31B run to back it, and the RX 9060 XT's 320 GB/s caps it lower. If you need 4-bit or better, no 16GB card is enough, so plan for 24GB.
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported. VRAM totals and throughput ceilings are calculated from the cited model config and file sizes and are labelled as calculations. Where no public run exists, the tables say so.
🛒 Products mentioned in this article
Amazon & eBay listings, plus full specs and alternatives on each product page.
As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.
Only at roughly 3-bit quantization. A 31-billion-parameter model at Q4_K_M uses about 4.8 bits per weight, which puts the weights alone near 18-19GB, more than a 16GB card holds. Q3_K_M and IQ3-class quants come in under 16GB but leave little headroom for the KV cache, so long contexts still push layers into system RAM.
Is Q3 on a 16GB card better than Q4 with CPU offload?
For interactive chat it usually is. Each layer offloaded to system RAM is read over PCIe and DDR memory, which are far slower than on-card GDDR, so generation speed drops sharply once even a few layers spill. Q4 with offload keeps more quality but answers more slowly. If you are doing batch summarization and are not waiting on the output, Q4 with offload is a reasonable choice.
Which backend should I use on the RX 9060 XT and Arc A770?
For the RX 9060 XT, llama.cpp runs through either ROCm or Vulkan. Vulkan is the easier install on consumer Windows machines, while ROCm on a supported Linux distro is often faster at prompt processing. For the Arc A770, the Vulkan and SYCL backends are the main options now that Intel has ended IPEX-LLM development. Benchmark both backends on your own driver version before you settle on one.
Should I sell my RTX 3060 12GB to step up to a 16GB card for Gemma 4 31B?
Only if you use the 31B model every day. On 12GB, Gemma 4 31B needs either a 2-bit quant or heavy offload, and both cost either quality or speed. If you mostly run 8B-14B models, the RTX 3060 12GB still handles them comfortably, and 16GB adds little for that workload.
How much does context length change the VRAM budget?
A lot. The KV cache grows linearly with context, so a quant that fits at 4K-8K context can overflow 16GB at 32K and beyond. Quantizing the KV cache to q8_0 roughly halves its footprint with a small quality cost, and flash attention reduces overhead further. The article's context table gives the sourced per-token figures for each setting.