Skip to main content

Gemma 4 31B on a 16GB GPU: RTX 5060 Ti vs RX 9060 XT vs Arc A770 Quant Table (2026)

Q3_K_M is the largest Gemma 4 31B quant that stays on a 16GB card. Here is how the RTX 5060 Ti, RX 9060 XT and Arc A770 compare on it.

Can a 16GB GPU run Gemma 4 31B? Per-quant VRAM table, KV-cache math and sourced tok/s for the RTX 5060 Ti, RX 9060 XT and Arc A770 vs a 12GB RTX 3060.

Quick Answer

Yes, a 16GB GPU can run Gemma 4 31B fully on the card, but only at 3-bit quants. Unsloth's Q3_K_M file is 14.74 GB, and on an RTX 5060 Ti 16GB it peaked at 15.0 GB of VRAM with an 8K q4_0 KV cache, per njannasch.dev's Gemma 4 on a 5060 Ti test. Q4_K_M (18.32 GB on the unsloth GGUF listing) does not fit. Pick the ASUS Dual RTX 5060 Ti 16GB. Its 448 GB/s of bandwidth and the CUDA backend give it the only published 16GB run, at 26–29 tok/s on UD-IQ3_XXS.

This page is for owners of 16GB cards, and buyers considering one, who have already tried Gemma 4 31B on 12GB and hit the offload wall. The 12GB RTX 3060 Gemma 4 31B guide covers the 2-bit and heavy-offload end of the range. This page covers the next tier up. Three 16GB cards are worth considering in 2026: NVIDIA's RTX 5060 Ti 16GB, AMD's RX 9060 XT 16GB and Intel's Arc A770 16GB. The RTX 3060 12GB is included as the baseline most readers are upgrading from. The format follows the Qwen 3.6 27B 16GB quant table, but Gemma 4 31B is a larger, differently built model, so the conclusions differ. Every VRAM figure below is either a published measurement or calculated from the model's own config, and each one says which.

Gemma 4 31B on a 16GB GPU: RTX 5060 Ti vs RX 9060 XT vs Arc A770 Quant Table (2026)

Small-model (7–9B) reference throughput

These figures are for 7–9B models, not the 30–35B models this article covers. The table holds model size fixed so the cards can be compared with each other; for the size in the title, use the article's own figures.

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead.

GPUVRAM Llama-3-8B class, Q4Street price Sources
NVIDIA GeForce RTX 5060 Ti 16 GB 71 tok/s11 runs · 8 sources $786street, all listings SpecPicks median of 11 runs; sources: LocalScore.ai (Mozilla Builders), RunAIHome, Runyard.dev, ComputingForGeeks +4 more
Radeon RX 9060 XT 16GB 16 GB 70.2 tok/s5 runs · 2 sources $529street, all listings SpecPicks median of 5 runs; sources: llama.cpp GitHub ROCm perf…, llama.cpp community benchmark…
NVIDIA GeForce RTX 3060 12 GB 55 tok/s23 runs · 9 sources $329MSRP SpecPicks median of 23 runs; sources: TYO Lab blog, Hardware Corner, llama.cpp GitHub (CUDA performanc…, Ajit Singh / Hardware-Corner +5 more
Intel Arc A770 16 GB 52.6 tok/s11 runs · 4 sources $349MSRP SpecPicks median of 11 runs; sources: llama.cpp GitHub Discussions, diunox/arc-inference-bench…, KnightLi llama.cpp GPU Scoreboard, Local AI Master

Which models fit on a RTX 5060 Ti?

The 30-35B class this article is about needs about 19.9 GB for its Q4 weights; on the RTX 5060 Ti, the weights do not fit, so layers spill to system RAM and PCIe bandwidth sets the speed. RTX 5060 Ti carries 16 GB of VRAM. The weights column is the range of real Q4_K_M files on Hugging Face for the models in each size class (about 0.6 GB per billion parameters), and the runtime plus a usable context window wants about 2 GB on top; every tokens-per-second figure is a median over community-reported Q4 runs SpecPicks tracks for this card, with the run count and the source beside it.

Showing the model sizes this article covers and the band either side. Every size from 3B to 70B+, for every card SpecPicks tracks, is in the local-LLM GPU table.

Model size Weights at Q4 Fits in 16 GB? Measured Left for context Source
20-27B (Gemma 3 27B, Mistral Small)A 27B Q4_K_M file (about 16.7 GB) is more than a 16 GB card holds, so 16 GB means a Q3 quant or partial CPU offload; a 24B loads with a short context. 20 GB holds the class with a usable context window, 24 GB with a long one. about 16.7 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —
30-35B (Qwen 3 32B, QwQ 32B)The step change. A 24 GB card holds this entirely in VRAM; below that it is CPU offload. about 19.9 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —
70B+ (Llama 3.3 70B, Qwen 2.5 72B)One 48 GB card or two 24 GB cards. A 32 GB card runs it only with layers in system RAM. about 42.5 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —

Every RTX 5060 Ti benchmark run, with its source → GPU picks for local LLM — the same table across every card we track How we source these numbers

Key Takeaways

  • The largest quant that runs fully on 16GB is Q3_K_M (14.74 GB), and only with a quantized KV cache. The published 5060 Ti run peaked at 15.0 GB at 8K context with q4_0 KV (njannasch.dev).
  • UD-IQ3_XXS (11.84 GB) is the comfortable 16GB quant. It generated 26–29 tok/s on an RTX 5060 Ti and held 65K context with q4_0 KV, per the same test.
  • Offload is a cliff, not a slope. A Q6_K file (23.4 GB) on the same card measured 0.62 tok/s generation in llama.cpp issue #26674.
  • Gemma 4's sliding-window layers keep long context cheap. 50 of the 60 layers attend to only 1,024 tokens, per the Gemma 4 31B config. Most of the KV cache stays a fixed size no matter how long the context gets.
  • Bandwidth sets the speed ceiling: 560 GB/s on the Arc A770, 448 GB/s on the RTX 5060 Ti and 320 GB/s on the RX 9060 XT (TechPowerUp RTX 5060 Ti, TechPowerUp RX 9060 XT, TechPowerUp Arc A770). Software maturity decides how much of that ceiling each card reaches.

Gemma 4 31B per-quant table on 16GB: what fits?

File sizes come from the unsloth/gemma-4-31B-it-GGUF listing, except Q2_K, which comes from bartowski/google_gemma-4-31B-it-GGUF because unsloth does not publish a plain Q2_K. BF16 comes from ggml-org/gemma-4-31B-it-GGUF.

The VRAM column is calculated. It adds the file size in GiB, the 8K f16 KV cache derived from the model config (1.8 GiB, explained in the KV section below), and about 0.5 GiB of llama.cpp compute buffer. The fit test is against roughly 15.5 GiB, the usable memory per RTX 5060 Ti 16GB after the driver reserve, as reported in LLMKube's dual-5060 Ti bake-off. This method lands within 0.1–0.3 GB of the two measured peaks in the njannasch.dev test. "Layers offloaded" is the excess divided by an average per-layer size (file ÷ 60 layers), rounded up.

QuantGGUF file sizeEst. VRAM at 8K (f16 KV)Fits fully in 16GB at 8K?Layers offloaded (of 60)Quality note
UD-IQ3_XXS11.84 GB13.3 GiB (measured peak 13.2 GB)Yes, room for 32K–65K with q4_0 KV0Lowest 3-bit; the speed pick
Q2_K12.63 GB14.1 GiBYes0Larger than IQ3_XXS but lower quality; skip it
Q3_K_S13.21 GB14.6 GiBYes0Safe 3-bit with f16 KV at 8K
Q3_K_M14.74 GB16.0 GiB (measured 15.0 GB with q4_0 KV)Partial with f16 KV; yes with q8_0/q4_0 KV0 with q4_0 KV, ~3 with f16 KVBest quality that stays on-card
IQ4_XS16.37 GB17.5 GiBNo~9First 4-bit option
Q4_K_M18.32 GB19.4 GiBNo~14The usual default; needs 24GB
Q5_K_M21.66 GB22.5 GiBNo~2124GB tier
Q6_K25.20 GB25.8 GiBNo (measured 0.62 tok/s offloaded)~2732GB tier
Q8_032.64 GB32.7 GiBNo~3448GB tier
BF1661.41 GB59.5 GiBNo~47Reference weights

Sources: file sizes from the unsloth, bartowski and ggml-org GGUF listings above. Measured peaks are from njannasch.dev. The 0.62 tok/s Q6_K figure is from llama.cpp issue #26674. Google also ships an official QAT build. The 17.65 GB Q4_0 file in google/gemma-4-31B-it-qat-q4_0-gguf is trained to hold quality at 4-bit, but it is still about 2 GB too large for a 16GB card.

Two caveats. First, loading the vision projector (mmproj, 1.2 GB) for image input adds about 1.1 GiB to every row, and that alone drops Q3_K_M from "fits" to "offload." Second, if the same card drives your desktop, subtract a few hundred MB from the 15.5 GiB budget.

Step 0: is 16GB actually your bottleneck?

Before you buy a card, check which constraint you are actually hitting. Answer three questions.

1. Are you offloading now, and by how much? If your current setup runs Gemma 4 31B Q4_K_M with 10+ layers in system RAM, a 16GB card does not fix that. Q4_K_M needs about 19.4 GiB at 8K, so it will still offload around 14 layers. The upgrade only helps if you are willing to drop to a 3-bit quant that fits completely.

2. Is Q3 on-card better than Q4 with offload? For interactive use, yes, by a wide margin. Every offloaded layer is read from system RAM over PCIe on every token. The issue #26674 report shows how bad that gets: 23.4 GB of Q6_K on a 16GB 5060 Ti generated 0.62 tok/s while prompt processing ran at 41.84 tok/s. The fully resident Q3_K_M ran at 17–19 tok/s on the same card model. Q4 with offload is reasonable only for batch jobs nobody is waiting on.

3. Is a 12GB RTX 3060 with offload good enough? On 12GB, the largest on-card option is about UD-IQ2_M (10.75 GB) with q4_0 KV. The calculated total of about 11.0 GiB leaves almost no room for context. If 31B is something you run occasionally and you can tolerate 2-bit quality, keep the 3060. If 31B is your daily model, the 3-bit tier is a real quality step and needs 16GB.

Worked example A: daily coding chat at 8K. Q3_K_M with --cache-type-k q4_0 --cache-type-v q4_0 and flash attention fits a headless 16GB card. Expect the 17–19 tok/s the njannasch.dev run measured on a 5060 Ti.

Worked example B: long-document Q&A at 64K. Drop to UD-IQ3_XXS with q4_0 KV. The same source held 65K context at a 13.8 GB peak and still generated 18 tok/s at 50K tokens.

Worked example C: you need 4-bit quality. No 16GB card does this without offload. Look at 24GB instead, or use Gemma 4's 26B MoE sibling, which the njannasch.dev test ran 3.5x faster on the same card.

How much VRAM does the KV cache add at 8K, 32K and 128K context?

Gemma 4 31B's config.json defines 60 layers. 50 are sliding-window layers (16 KV heads, head size 256, 1,024-token window) and 10 are full-attention layers (4 KV heads, head size 512). The two types use memory very differently:

  • Sliding layers: 800 KiB per token in f16, but llama.cpp sizes this cache to the window plus one micro-batch (1,024 + 512 = 1,536 cells by default). That is a fixed ~1.17 GiB in f16, no matter how long the context is.
  • Full-attention layers: 80 KiB per token in f16, growing linearly with context.
ContextFixed sliding cache (f16 / q8_0 / q4_0)Growing global cache (f16 / q8_0 / q4_0)Total KV (f16 / q8_0 / q4_0)
8K1.17 / 0.62 / 0.33 GiB0.62 / 0.33 / 0.18 GiB1.80 / 0.95 / 0.51 GiB
32K1.17 / 0.62 / 0.33 GiB2.50 / 1.33 / 0.70 GiB3.67 / 1.95 / 1.03 GiB
128K1.17 / 0.62 / 0.33 GiB10.0 / 5.31 / 2.81 GiB11.2 / 5.94 / 3.14 GiB

These figures are calculated from config.json, assuming llama.cpp stores K and V separately for every layer. The config's attention_k_eq_v flag could let a runtime halve the global portion, but the measured peaks above are consistent with the conservative figure.

What this means on 16GB:

  • Do not pass --swa-full. That flag makes the sliding layers cache the full context too. At 32K that is 25 GiB of f16 cache, and it is the most common reason Gemma 4 runs out of memory on 16GB cards, per njannasch.dev.
  • q8_0 KV is the default to reach for. It roughly halves the cache, which is the margin that lets Q3_K_M stay on-card. q4_0 saves more and is what the 65K run used, but it costs some long-context recall quality.
  • 128K is out of reach for the dense 31B on 16GB. Even IQ3_XXS with q4_0 KV comes to about 14.7 GiB in this model. The njannasch.dev 131K configuration loaded at a 15.5 GB peak but ran out of memory before reaching 108K tokens of generation.

Spec-delta table: RTX 5060 Ti 16GB vs RX 9060 XT 16GB vs Arc A770 16GB vs RTX 3060 12GB

CardVRAMMemory bandwidthLaunch MSRPBoard powerllama.cpp backend
GeForce RTX 5060 Ti 16GB16GB GDDR7, 128-bit448 GB/s$429180WCUDA
Radeon RX 9060 XT 16GB16GB GDDR6, 128-bit320 GB/s$349160WROCm (HIP) or Vulkan
Arc A770 16GB16GB GDDR6, 256-bit560 GB/s$349225WSYCL or Vulkan
GeForce RTX 3060 12GB12GB GDDR6, 192-bit360 GB/s$329170WCUDA

Specs and launch prices are from TechPowerUp's database entries: RTX 5060 Ti 16 GB, RX 9060 XT 16 GB, Arc A770 and RTX 3060 12 GB. These are launch prices, all at MSRP. Street prices move weekly; the live price buttons on this page carry today's numbers.

All three 16GB cards hold exactly the same quants, so the per-quant table applies to each. They differ in how fast they read those weights. On paper, the Sparkle Arc A770 ROC OC 16GB has the widest bus of the four, 256-bit at 560 GB/s, which is 25% more than the 5060 Ti and 75% more than the 9060 XT. Whether you get that bandwidth in practice depends on the backend, as the next section shows.

Which 16GB card generates Gemma 4 31B tokens fastest?

As of October 2026, only one card has a published single-card Gemma 4 31B run that fits on the card: the RTX 5060 Ti. For the others, the table gives the bandwidth ceiling. The ceiling is memory bandwidth divided by file size, a physical upper bound rather than a measurement.

CardUD-IQ3_XXS (11.84 GB) gen tok/sQ3_K_M (14.74 GB) gen tok/sBackendSource
RTX 5060 Ti 16GB26–29 measured (ceiling 37.8)17–19 measured, q4_0 KV (ceiling 30.4)CUDAnjannasch.dev
RX 9060 XT 16GBno public run (ceiling 27.0)no public run (ceiling 21.7)varies by backendceiling: TechPowerUp bandwidth ÷ unsloth file size
Arc A770 16GBno public run (ceiling 47.3)no public run (ceiling 38.0)varies by backendceiling: TechPowerUp bandwidth ÷ unsloth file size
RTX 3060 12GBdoes not fit at 8Kdoes not fitCUDAsee per-quant table

The 5060 Ti reached about 77% of its ceiling on IQ3_XXS and about 62% on Q3_K_M with q4_0 KV, where dequantizing the cache adds work. If the RX 9060 XT reached the same 77%, it would land around 21 tok/s on IQ3_XXS. Treat that as a projection until someone publishes the run.

To compare how much of their bandwidth each card actually delivers in real software, here is the median of each card's published 7–9B generation runs (llama.cpp and Ollama) from the SpecPicks benchmark pages:

Card7–9B gen medianRuns / sourcesBenchmark page
RTX 5060 Ti 16GB59.2 tok/smedian of 7 published runs (5 sources)/benchmarks/nvidia-geforce-rtx-5060-ti-16gb
RX 9060 XT 16GB67.6 tok/smedian of 3 published runs (3 sources)/benchmarks/amd-radeon-rx-9060-xt-16gb
Arc A770 16GB37.0 tok/smedian of 8 published runs (4 sources)/benchmarks/intel-arc-a770
RTX 3060 12GB57.4 tok/smedian of 16 published runs (8 sources)/benchmarks/nvidia-geforce-rtx-3060-12-gb

Read these medians with care. Two of the three RX 9060 XT runs use Llama 2 7B q4_0, a smaller file than the Llama 3.1 8B Q4_K_M behind most of the NVIDIA runs, so the 9060 XT median is flattered. The A770's published runs span 15–67 tok/s depending on backend and build. In the llama.cpp Vulkan scoreboard (#10879), the A770 posts 52.6 tok/s on Llama 2 7B q4_0, against 70.5 for the RX 9060 XT and 75.9 for the RTX 3060 on the same harness. The card with the most bandwidth on paper finishes last on the same benchmark. Scaling to 31B does not change that pattern, because generation stays bandwidth-bound and backend efficiency carries over.

Prefill vs generation: why prompt processing favours the RTX 5060 Ti's CUDA path

Generation reads every weight once per token, so it is limited by bandwidth. Prefill processes the whole prompt in parallel, so it is limited by compute and kernel quality. It decides how long you wait for the first token, and with a 31B model and long pastes, that wait is the delay you notice.

CardModel / quantPrefill tok/sBackendSource
RTX 5060 Ti 16GBLlama 3.1 8B Q4_K_M2,361CUDA (llamafile)LocalScore RTX 5060 Ti
RTX 3060 12GBLlama 3.1 8B Q4_K_M1,490CUDA (llamafile)LocalScore RTX 3060
RX 9060 XT 16GBLlama 2 7B q4_02,142Vulkanllama.cpp #10879
RTX 3060 12GBLlama 2 7B q4_01,816Vulkanllama.cpp #10879
Arc A770 16GBLlama 2 7B q4_01,074Vulkanllama.cpp #10879

On the same LocalScore harness, the 5060 Ti prefills 58% faster than the 3060. On the Vulkan scoreboard, the 9060 XT prefills 18% faster than the 3060, and the A770 runs at about half the 9060 XT's rate. Chaining the two ratios through the RTX 3060 puts the 5060 Ti ahead of the 9060 XT on prefill. That is indirect, since the harnesses differ, but it lines up with Blackwell's 5th-generation tensor cores and the maturity of the CUDA flash-attention kernels.

Gemma 4's architecture also helps prefill. The 50 sliding layers attend to only 1,024 tokens, so attention cost on those layers stops growing with prompt length. The 10 global layers carry the long-context cost.

Backend gotchas per card

RTX 5060 Ti 16GB (CUDA). Use a llama.cpp build compiled for compute capability 12.0 (Blackwell). Older prebuilt binaries fall back to PTX JIT and start slowly. The most-missed step: check that -ngl 99 actually put every layer on the GPU. In issue #26674, the reporter saw 100% CPU and under 15% GPU utilization during Gemma 4 generation. A maintainer attributed the slow speed to an oversized dense model spilling into system RAM. If nvidia-smi shows less VRAM in use than your file size, layers are on the CPU.

RX 9060 XT 16GB (ROCm or Vulkan). Vulkan is the easier install and works on Windows. ROCm on Linux needs a release that lists RDNA 4 (gfx1200) as supported. The most-missed step: on ROCm, set HSA_OVERRIDE_GFX_VERSION only if your ROCm build predates official gfx1200 support. On a current build it does more harm than good. The parsapp RX 9060 XT benchmark repo found Vulkan slightly ahead on generation at 14B, so start there.

Arc A770 16GB (SYCL or Vulkan). Intel's IPEX-LLM repository was archived in January 2026, so the guides that start with "install IPEX-LLM" are out of date. Use llama.cpp's own Vulkan or SYCL backend. The most-missed step: try Vulkan first. In llama.cpp issue #19918, an A770 owner measured SYCL at 10 tok/s generation versus 68 on Vulkan with the same MoE models. Also enable Resizable BAR in the BIOS. Arc cards lose significant performance without it.

Common pitfalls on all three cards:

  1. Leaving the vision projector loaded. --mmproj adds about 1.2 GB. Load it only when you send images.
  2. Using --swa-full from an old guide. It undoes Gemma 4's cache savings (see the KV section).
  3. Running an outdated build. Gemma 4 support, including its chat template and iSWA cache, needs a 2026 llama.cpp build.
  4. Benchmarking with the browser open. Hardware-accelerated browsers can hold several hundred MB of VRAM, and on 16GB that margin is what Q3_K_M needs.

Perf-per-dollar and perf-per-watt math (at MSRP)

All prices are launch MSRP from the spec table, so there is one consistent price basis. Power is rated board power, not measured draw. Two metrics are shown: the measured 7–9B median, and the calculated Gemma 4 31B IQ3_XXS bandwidth ceiling.

Card7–9B median tok/s per $100 MSRP7–9B median tok/s per 100W31B IQ3_XXS ceiling per $100 MSRP31B IQ3_XXS ceiling per 100W
RTX 5060 Ti 16GB13.832.98.821.0
RX 9060 XT 16GB19.442.27.716.9
Arc A770 16GB10.616.413.621.0
RTX 3060 12GB17.433.7n/a (does not fit)n/a

The two metrics point in different directions. On measured small-model throughput, the RX 9060 XT gives the most tokens per launch dollar and per watt, although its median is flattered by smaller test files. On the bandwidth ceiling for this specific 31B model, the A770 leads per dollar, and the A770 and 5060 Ti tie per watt. In practice, the 5060 Ti is the only card with a measured 31B result (26–29 tok/s), and that measurement counts for more than either ceiling.

Verdict matrix

Get the RTX 5060 Ti 16GB if…

  • Gemma 4 31B is your daily model. It is the only 16GB card with a published on-card run (26–29 tok/s on IQ3_XXS, 17–19 tok/s on Q3_K_M).
  • You paste long prompts and want the fastest prefill of the group.
  • You want CUDA's ecosystem: llama.cpp, vLLM, ExLlama and Ollama all support Blackwell.

Get the RX 9060 XT 16GB if…

  • Most of your work is 7–14B models, where its measured throughput matches or beats the 5060 Ti for less money at launch.
  • You accept roughly 29% less bandwidth on 31B (320 vs 448 GB/s) in exchange for 160W board power.
  • You are comfortable with llama.cpp's Vulkan build. See the ASUS Dual Radeon RX 9060 XT 16GB.

Get the Arc A770 16GB if…

  • You already own one, or find it well below the other two. Its 560 GB/s is the highest ceiling here, but published llama.cpp runs reach well below it.
  • You will use the Vulkan backend and keep Resizable BAR on.

Stay on the RTX 3060 12GB with offload if…

Buy the ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC and run Gemma 4 31B Q3_K_M with q8_0 or q4_0 KV cache at 8K. Switch to UD-IQ3_XXS with q4_0 KV when you need 32K–65K context. The deciding spec is memory bandwidth backed by mature software. The 5060 Ti's 448 GB/s delivered a measured 77% of its ceiling on this exact model, while the A770's larger paper ceiling has no 31B run to back it, and the RX 9060 XT's 320 GB/s caps it lower. If you need 4-bit or better, no 16GB card is enough, so plan for 24GB.

Live price comparison

See current prices for the two main contenders side by side at RTX 5060 Ti 16GB vs RX 9060 XT 16GB.

Citations and sources

All sources accessed 2026-10-10.

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported. VRAM totals and throughput ceilings are calculated from the cited model config and file sizes and are labelled as calculations. Where no public run exists, the tables say so.

Products mentioned in this article

Amazon & eBay listings, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

Please Buy Intel GPUs. - Arc A750 & A770 Review — Linus Tech Tips on YouTube

Frequently asked questions

Does Gemma 4 31B fit entirely in 16GB of VRAM?
Only at roughly 3-bit quantization. A 31-billion-parameter model at Q4_K_M uses about 4.8 bits per weight, which puts the weights alone near 18-19GB, more than a 16GB card holds. Q3_K_M and IQ3-class quants come in under 16GB but leave little headroom for the KV cache, so long contexts still push layers into system RAM.
Is Q3 on a 16GB card better than Q4 with CPU offload?
For interactive chat it usually is. Each layer offloaded to system RAM is read over PCIe and DDR memory, which are far slower than on-card GDDR, so generation speed drops sharply once even a few layers spill. Q4 with offload keeps more quality but answers more slowly. If you are doing batch summarization and are not waiting on the output, Q4 with offload is a reasonable choice.
Which backend should I use on the RX 9060 XT and Arc A770?
For the RX 9060 XT, llama.cpp runs through either ROCm or Vulkan. Vulkan is the easier install on consumer Windows machines, while ROCm on a supported Linux distro is often faster at prompt processing. For the Arc A770, the Vulkan and SYCL backends are the main options now that Intel has ended IPEX-LLM development. Benchmark both backends on your own driver version before you settle on one.
Should I sell my RTX 3060 12GB to step up to a 16GB card for Gemma 4 31B?
Only if you use the 31B model every day. On 12GB, Gemma 4 31B needs either a 2-bit quant or heavy offload, and both cost either quality or speed. If you mostly run 8B-14B models, the RTX 3060 12GB still handles them comfortably, and 16GB adds little for that workload.
How much does context length change the VRAM budget?
A lot. The KV cache grows linearly with context, so a quant that fits at 4K-8K context can overflow 16GB at 32K and beyond. Quantizing the KV cache to q8_0 roughly halves its footprint with a small quality cost, and flash attention reduces overhead further. The article's context table gives the sourced per-token figures for each setting.

— Mike Perry · Updated 2026-10-10

Parts this article names

Amazon Associate — prices tracked 2026-10-10, may vary.