Gemma 4 punches well above its weight class on instruction-following and conversational quality benchmarks. The 9B variant is the best "fits in 12 GB VRAM and feels like a much bigger model" pick in the 2023 open-weight space — perfect for a daily-driver chat rig on consumer hardware.
Hardware tiers for running Gemma 4
🧪
Minimum (2B / 7B at Q4)
Gemma 4 2B at Q4 fits even on integrated graphics. 7B Q4 needs 8 GB; 12 GB lets you keep full 128K context active.
VRAM8 GB
System RAM16 GB DDR5
Throughput~140 tok/s on 2B Q4 · ~75 tok/s on 7B Q4 (RTX 4060)
For chat / writing quality at small sizes, Gemma 4 9B beats Llama 4's 8B variant on most leaderboards. For coding or tool-use, Llama 4 / Qwen3 are stronger. For pure conversational quality on a 12-16 GB card, Gemma 4 9B is the choice.
Is Gemma 4 multimodal?
The 9B and 27B variants have native vision support in v4.1. The 2B and 7B base variants are text-only.
What's the cheapest GPU for Gemma 4 27B?
Used RTX 3090 ($800-1100). 24 GB VRAM holds 27B at Q4 with comfortable context. RTX 4090 / 5090 give better tok/s but the 3090 is the price/perf winner.
More guides & deep dives from the SpecPicks archive