Skip to main content
Updated 2026-09-16 3 hardware tiers Quant: Q4_K_M Context: 128K

Run Qwen3 Locally — Hardware Tiers, Tok/s & Build Guide

Alibaba's flagship open-weight model family with strong code, math, and tool-use performance. Available variants: 1.7B 4B 7B 14B 32B 72B.

As an Amazon Associate, SpecPicks earns from qualifying purchases. See our review methodology.

Why Qwen3 matters in 2026

Qwen3 currently leads the open-weight leaderboard on tool-use and code benchmarks at every parameter tier. The 32B variant is the sweet-spot model for a single 24-32GB GPU local rig in 2023 — runs cleanly at Q4_K_M on an RTX 4090 or RTX 5090, hits ~60 tok/s on 5090 with 128K context active.

Hardware tiers for running Qwen3

Recommended (32B at Q4)

Sweet spot for daily-driver Qwen3 use. 24 GB fits 32B at Q4_K_M with 32–64K context comfortably; 32 GB (RTX 5090) extends context to full 128K and keeps latency low.

VRAM24 GB
System RAM64 GB DDR5
Throughput~25 tok/s on 32B Q4 (RTX 4090) · ~38 tok/s (RTX 5090 with 32B at higher quant)
ASUS ROG Strix GeForce RTX 4090 OC Edition Gaming Graphics Card (PCIe 4.0, 24GB GDDR6X, HDMI 2.1a, DisplayPort 1.4a), 3 Year Warranty

ASUS ROG Strix GeForce RTX 4090 OC Edition Gaming Graphics Card (PCIe 4.0…

$4449.99

*Price sourced from Amazon.com. Price and availability subject to change.

Luxury (72B at Q4)

Dual RTX 3090 (used) is the best price/perf 48 GB tier in 2023 — pool VRAM via NVLink, run 72B at Q4 with comfortable context. RTX PRO 6000 Blackwell (96 GB) is the single-card alternative if budget allows.

VRAM48 GB
System RAM128 GB DDR5
Throughput~14 tok/s on 72B Q4 (dual RTX 3090 NVLink)
Fasgear 16pin GPU Cable to 4x8 Pin Pcie Extension, Black | PCI-e 5.1 12VHPWR Extender Cord, 40cm Sleeved, 12V-2x6 Cable, For GeForce RTX 5090 5080 5070 5070Ti 4090 4080 4070ti 3090Ti

Fasgear 16pin GPU Cable to 4x8 Pin Pcie Extension, Black | PCI-e 5.1 12VHPWR…

$29.99

*Price sourced from Amazon.com. Price and availability subject to change.

Related Comparisons

Frequently Asked Questions

What's the smallest GPU that runs Qwen3 32B?

A 24 GB card (RTX 4090, RTX 3090, or RTX 5090) at Q4_K_M quantization. With 16 GB cards you can fit 14B comfortably but the 32B variant requires either Q3 quant (quality drop) or CPU offload (latency drop).

Q4_K_M vs FP16 — what's the quality cost?

On benchmarks, Q4_K_M loses about 1-2% on average vs FP16 — well below practical-use noticeable. Q3_K_M starts to show: 4-6% drop and noticeable on edge cases. Always pick Q4_K_M as the default if VRAM permits.

Can I run Qwen3 on a Mac?

Yes, with caveats. Mac Studio M3 Ultra (192 GB unified) runs 72B at Q4 well — slower than dual RTX 3090 (~10 tok/s vs ~14) but silent, low-power, and capacity is enormous. M4 Max with 128 GB also works for up to 32B comfortably.

Is Qwen3 better than Llama 4 / DeepSeek V3?

Depends on your workload. Qwen3 leads on tool-use and coding benchmarks; Llama 4 has better creative writing; DeepSeek V3 has the best math and reasoning at very large sizes. For a local-rig daily driver, Qwen3 32B is the most balanced.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →

— SpecPicks Editorial · Last verified 2026-09-16