Skip to main content
Best GPUs for Mixtral 8x22B Local Inference in 2026

Best GPUs for Mixtral 8x22B Local Inference in 2026

Mixtral 8x22B needs 86 GB at Q4. These are the GPU and RAM builds that actually run it, from a $440 offload box to four RTX 3090s.

Best GPUs for Mixtral 8x22B in 2026: VRAM and RAM needs per quant, llama.cpp expert offload, and five picks from RTX 3060 budget builds to multi-3090 rigs.

Hardware at a Glance

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Street price is the second-lowest tracked listing within a sane band of MSRP, so no single listing sets it; prices move daily.

GPUVRAM Llama-3-8B class, Q4Street price Benchmark source
NVIDIA GeForce RTX 5090 32 GB 185.9 tok/s4 runs · 3 sources $1,999MSRP Hardware Corner
NVIDIA GeForce RTX 4090 24 GB 126.4 tok/s8 runs · 7 sources $3,149street, all listings Hardware Corner
NVIDIA GeForce RTX 3090 24 GB 93.9 tok/s6 runs · 4 sources $1,780street, all listings MyAIHardware
NVIDIA GeForce RTX 5060 Ti 16 GB 65.1 tok/s18 runs · 11 sources $429MSRP GPU Battle
NVIDIA GeForce RTX 3060 12 GB 57.4 tok/s30 runs · 16 sources $392street, all listings smeltcore.com

As an Amazon Associate, SpecPicks earns from qualifying purchases. See our review methodology.

Quick Answer

No single consumer GPU holds Mixtral 8x22B. It has 141B total parameters and 39B active per token, per Mistral's release post, and the Q4_K_M GGUF is 85.6 GB. The best single-card setup is an RTX 4090 with 128 GB of DDR5 and the expert layers offloaded to the CPU, at roughly 3–4 tokens/sec. For interactive speed, you need four RTX 3090s.

Step 0: full VRAM or expert offload?

Before you pick a GPU, decide which of two fundamentally different builds you want.

Full-VRAM builds put every weight on GPUs. At Q4_K_M that means about 86 GB of weights plus KV cache, so four 24 GB cards. At Q3 it means three 24 GB cards or two 32 GB cards. These builds give interactive generation speed, in the teens of tokens per second, but they need a workstation board, a 1,600 W+ power budget and serious airflow.

Expert-offload builds keep the attention and shared layers on one GPU and put most of the 8 experts' feed-forward weights in system RAM. This works much better for Mixtral than for a dense 70B model, because only 2 of 8 experts fire per token. The CPU reads about a quarter of the expert weights for each token, not all of them. llama.cpp supports this directly with --n-cpu-moe and --override-tensor. In this mode, generation speed is set mostly by system-RAM bandwidth, not by the GPU. The GPU decides how many experts live in fast VRAM, and it handles prompt processing.

The rest of this guide covers both builds. The short version: one RTX 4090 plus 128 GB of DDR5 is the best single-GPU experience. A 16 GB RTX 5060 Ti gets within about 20% of it in generation for a sixth of the price. Four RTX 3090s are the only consumer route to genuinely interactive speed.

VRAM and quantization requirements

File sizes are from bartowski's Mixtral-8x22B GGUF repository. The KV cache adds about 0.22 MB per token at FP16: 1.8 GB at 8K, 7.3 GB at 32K and 14.7 GB at the full 64K context.

QuantFile sizeFull-VRAM needsOffload build needsWhich pick handles it
Q2_K52.1 GB2× 32 GB (64 GB)12–16 GB GPU + 64 GB RAMRTX 5090 pair, or any pick with offload
IQ3_XS58.2 GB3× 24 GB or 2× 32 GB12–16 GB GPU + 64–96 GB RAM3× RTX 3090, or offload
Q3_K_M67.8 GB3× 24 GB (tight, ~4K ctx)16–24 GB GPU + 96 GB RAM3× RTX 3090, or offload
Q4_K_M85.6 GB4× 24 GB (96 GB)24–32 GB GPU + 96–128 GB RAM4× RTX 3090, or RTX 4090/5090 offload
Q5_K_M100.0 GB5× 24 GB24–32 GB GPU + 128 GB RAMRTX 5090 offload only
Q8_0~150 GB7× 24 GB32 GB GPU + 192 GB RAMWorkstation territory

Most people should run Q4_K_M. Q3 quants cost measurable quality on a model whose main appeal is quality. Q5 and above rarely justify doubling the RAM bill.

Comparison at a glance

PickBest ForKey SpecPrice Range (Amazon, Sept 24, 2026)Verdict
MSI RTX 4090 24GBBest overall single card24 GB, 1,008 GB/s$2,949.99Most experts in VRAM at a sane price, fastest single-card prefill
GIGABYTE RTX 5060 Ti 16GBest value16 GB GDDR7, 448 GB/s$439.9980% of the 4090's offload generation for 15% of the price
ASUS TUF RTX 3090 24GBBest for multi-GPU24 GB, 936 GB/s$1,999.99 new (buy used)Four of them run Q4 fully in VRAM
ASUS TUF RTX 5090 32GBBest performance32 GB GDDR7, 1,792 GB/s$6,699.99Most layers on one card, but priced far above MSRP
MSI RTX 3060 12GBBudget pick12 GB, 360 GB/s$479.99It works for batch jobs, not for chat

🏆 Best Overall: MSI Gaming GeForce RTX 4090 24GB

The MSI Gaming GeForce RTX 4090 24GB (Suprim Liquid X) is the best single card for Mixtral 8x22B expert offload. At Q4_K_M with an 8K context, 24 GB holds the ~3.3 GB of shared attention and embedding weights and the KV cache, plus about 17 GB of expert tensors. That's roughly a fifth of all expert weights, so the CPU reads about 16 GB per token instead of 20.5 GB.

Estimated generation: 3–4 tok/s with dual-channel DDR5-6000 and 128 GB of RAM. This is our bandwidth-derived estimate, not a measured result. We couldn't find a reproducible public 8x22B benchmark with a matching setup, and every number we found mixed runtimes and RAM configurations. For reference, the same card runs the much smaller Qwen3 30B-A3B MoE fully in VRAM at 207 tok/s in Hardware Corner's llama.cpp runs. That shows how much offload costs compared with a model that fits.

✅ Most VRAM for the money after used 3090s, and the fastest single-card prefill of the sensible options ✅ FP8 tensor support if you later move to vLLM serving ✅ The liquid-cooled Suprim keeps the card quiet under sustained load ❌ In offload mode, generation is RAM-bound. Only about 20% separates it from the $440 5060 Ti ❌ 450 W TGP and an 850 W+ PSU recommendation ❌ The 240 mm radiator needs a case mount point

Prices change often. Check the current listing before you buy.

See Full Details →

💰 Best Value: GIGABYTE GeForce RTX 5060 Ti Gaming OC 16G

The GIGABYTE RTX 5060 Ti Gaming OC 16G is the card to buy if you want Mixtral 8x22B for batch work, and you'd rather spend the savings on RAM. Its 16 GB holds the shared layers, an 8K KV cache and about 9 GB of experts. That leaves the CPU reading around 18 GB per token, only about 12% more than on the 4090.

Estimated generation: 2.5–3.5 tok/s with the same 128 GB DDR5-6000 setup. The gap to the 4090 is small because system RAM, not the GPU, is the bottleneck. The gap in prompt processing is larger: the 5060 Ti has about a third of the 4090's tensor throughput, and its PCIe 5.0 x8 link limits how quickly llama.cpp can stream expert weights for large batches.

✅ $439.99 buys most of the offload experience ✅ 180 W TGP, easy to cool and to power ✅ Also a strong card for 14B models at 32K context, fully in VRAM ❌ Prefill on long prompts is noticeably slower than on the 4090 ❌ The 128-bit bus limits its speed on smaller models that fit entirely in VRAM

Prices change often. Check the current listing before you buy.

See Full Details →

🎯 Best for Multi-GPU: ASUS TUF Gaming RTX 3090 24GB

If you want Mixtral 8x22B at interactive speed, multiple ASUS TUF Gaming RTX 3090 OC 24GB cards are the consumer route. Four cards (96 GB) hold Q4_K_M fully in VRAM with room for an 8–16K context. Three cards (72 GB) hold IQ3_XS (58.2 GB) comfortably, or Q3_K_M at a short context.

Estimated generation: 15–20 tok/s for four 3090s at Q4_K_M with llama.cpp layer splitting. That's our estimate: each token reads about 24 GB of active weights at 936 GB/s per card, discounted for sequential layer-split overhead. Tensor-parallel serving in vLLM can push past that on a board with x8/x8/x8/x8 PCIe lanes.

✅ The only sub-workstation-GPU path to full-VRAM Q4 ✅ NVLink bridges exist for pairs, though llama.cpp doesn't need them ✅ The used market is deep, and eBay listings are far below the $1,999.99 new-stock price shown here ❌ Four cards at 350 W each need a 1,600 W+ PSU, or two PSUs, and a dedicated circuit ❌ At 2.7 slots, four TUF cards need risers or an open-frame chassis ❌ No FP8, and hot GDDR6X memory on the back of the board

Prices change often. Check the current listing before you buy.

See Full Details →

⚡ Best Performance: ASUS TUF Gaming RTX 5090 32GB

The ASUS TUF Gaming RTX 5090 32GB puts the most of Mixtral 8x22B on a single card. With 32 GB it holds about 25 GB of experts next to the shared layers, roughly 30% of the expert weights. Its 1,792 GB/s of bandwidth and Blackwell tensor cores give it the fastest prefill of any consumer card.

Estimated generation: 3.5–4.5 tok/s in single-card offload with 128 GB of DDR5. Two 5090s (64 GB) can hold IQ3_XS or Q2_K fully in VRAM, where we estimate 30+ tok/s. On a model this size, that's the fastest consumer configuration there is.

✅ The most expert weights in VRAM on one card ✅ The fastest prefill for long-context RAG over Mixtral's 64K window ✅ A pair runs Q2/IQ3 fully in VRAM ❌ The current $6,699.99 listing is more than three times the $1,999 MSRP ❌ 575 W TGP, a 3.6-slot cooler and a 1,000 W PSU recommendation ❌ In single-card offload it's only about 15% faster than the 4090 at generation

Prices change often. Check the current listing before you buy.

See Full Details →

🧪 Budget Pick: MSI Gaming GeForce RTX 3060 12GB

The MSI Gaming RTX 3060 12GB, or the shorter ZOTAC RTX 3060 Twin Edge OC 12GB, will run Mixtral 8x22B. Use --cpu-moe (all experts on the CPU) or --n-cpu-moe 53 (keep the last three layers' experts on the GPU) and the model loads, provided you have 96–128 GB of RAM. The GPU holds the shared layers and KV cache and does the attention math. The CPU does nearly all of the expert work.

Estimated generation: 2–3 tok/s with 128 GB of DDR5-6000, and lower on DDR4 platforms. That's fine for overnight summarization, batch classification or evaluation runs. It's too slow for interactive coding. Treat it as a way to try the model before you spend on a 24 GB card.

✅ The cheapest way to load a 141B model at all ✅ 170 W TGP, fits in almost any case ❌ Prompt processing on long inputs is slow, often minutes for a 16K prompt ❌ 12 GB leaves almost no room for experts, so generation is essentially CPU-bound

Prices change often. Check the current listing before you buy.

See Full Details →

What to look for in a GPU for Mixtral 8x22B

VRAM vs system RAM split

In offload mode, each extra GB of VRAM moves about 1.2% of the expert weights off the CPU's per-token read. Going from 16 GB to 24 GB saves roughly 2 GB of RAM traffic per token, which is worth about 10–15% more generation speed. VRAM still matters, but it matters less than it does for dense models. Size your system RAM first: file size, plus 8–10 GB for the OS, minus what fits in VRAM.

Memory bandwidth, both kinds

GPU bandwidth decides speed only in full-VRAM builds. In offload builds, dual-channel DDR5-6000 delivers about 96 GB/s in theory and 60–75 GB/s in practice, and that's your ceiling. A 2×64 GB DDR5 kit usually runs faster than a 4×32 GB kit, because four DIMMs often force consumer boards down to DDR5-4800 or lower. Check your motherboard's QVL.

MoE expert offload support in llama.cpp

llama.cpp has first-class MoE offload. -ngl 99 puts every layer on the GPU. Then --cpu-moe moves all expert tensors back to system RAM, or --n-cpu-moe N does it for the first N layers only. The regex form -ot "blk\.(\d+)\.ffn_.*_exps\.=CPU" gives finer control. Raising the micro-batch size (-ub 2048) speeds up prompt processing noticeably in offload mode. Ollama (mixtral:8x22b) handles the split automatically but gives you less control.

PSU and transient headroom

Budget for the whole system, not only the GPU's TGP. A 4090 box with a 16-core CPU under full load pulls 700–800 W, and brief transients go higher, so use an ATX 3.x 1,000 W unit. Four 3090s need around 1,600–2,000 W of PSU capacity. That is at or above what one 15 A, 120 V household circuit can safely supply.

Multi-GPU slot spacing

Consumer boards rarely space more than two x16 slots far enough apart for 2.7–3.6-slot coolers. For three or four cards, plan on an open-frame mining chassis with PCIe 4.0 risers, or on a workstation board (Threadripper or Xeon W) with four full-bandwidth slots.

Real-world example builds

  • Batch summarizer, about $1,000 in GPU + RAM: RTX 5060 Ti 16G, 128 GB DDR5-6000 (2×64 GB), Q4_K_M, --n-cpu-moe 50. Expect about 3 tok/s. It runs overnight jobs well.
  • Single-card workstation: RTX 4090, 128 GB DDR5-6000, Q4_K_M with 16K context. Expect about 3.5 tok/s generation with usable prefill. That's good for occasional interactive queries.
  • Four-card server: 4× used RTX 3090 on an open frame, Threadripper board, 2× 1,000 W PSUs, Q4_K_M fully in VRAM. Expect 15–20 tok/s. That's genuinely interactive.

When not to buy for Mixtral 8x22B

If you want an MoE model for daily interactive use on one consumer card, newer small-active-parameter MoEs fit far better than a 141B model. Qwen3 30B-A3B runs at over 200 tok/s fully inside a 24 GB card. Our best hardware for MoE LLMs guide compares the options. Choose Mixtral 8x22B for its Apache 2.0 license, its 64K context and a well-understood model you can pin, not because you want the fastest local chat.

FAQ

How much memory does Mixtral 8x22B need at Q4? Mixtral 8x22B has about 141B total parameters, so a Q4_K_M GGUF is roughly 80-86 GB before KV cache. That rules out any single consumer GPU for full-VRAM inference. Your options are three or four 24 GB cards, or a mix of GPU VRAM plus 96-128 GB of system RAM, with the expert layers offloaded to the CPU.

Why does expert offload work better for Mixtral than for a dense 70B model? Only two of the eight experts fire per token, so roughly 39B parameters are active, per Mistral's release post. When the expert tensors sit in system RAM, the CPU reads a fraction of the model per token. A dense 70B model has to read all of its weights every time. That makes MoE offload workable on a 12-24 GB card.

Is Mixtral 8x22B still worth running in 2026? Newer MoE releases often score higher on reasoning, but Mixtral 8x22B keeps its Apache 2.0 license, 64K context, and mature tooling support in llama.cpp, vLLM, and Ollama. If you need a permissively licensed model with well-understood behavior for production pipelines, it remains a reasonable pick. Otherwise, compare it with current open-weight MoE models first.

How much system RAM do I need alongside the GPU? For Q4 with expert offload, plan on 128 GB of system RAM if the GPU has 12-16 GB, or 96 GB with a 24-32 GB card. Dual-channel DDR5 bandwidth sets the ceiling on offloaded token speed, so faster rated kits help. Populating four DIMMs on consumer boards often forces lower memory speeds, so check your board's QVL.

Can I run it on a single RTX 3060 12GB at all? Yes, with llama.cpp's MoE CPU-offload options and enough system RAM, but expect single-digit tokens per second at best. The GPU holds attention and shared layers, and the CPU handles the experts. That's fine for batch summarization or overnight jobs, but too slow for interactive coding. Treat it as a starting point before a 24 GB upgrade.

Sources

  1. Mistral AI — Cheaper, Better, Faster, Stronger (Mixtral 8x22B release) (accessed September 24, 2026)
  2. Hugging Face — mistralai/Mixtral-8x22B-Instruct-v0.1 model card (accessed September 24, 2026)
  3. GitHub — ggml-org/llama.cpp (accessed September 24, 2026)
  4. Hugging Face — bartowski/Mixtral-8x22B-v0.1-GGUF quant sizes (accessed September 24, 2026)
  5. Hardware Corner — RTX 3090 / 4090 / 5090 llama.cpp benchmarks (accessed September 24, 2026)

Tokens-per-second figures for Mixtral 8x22B in this guide are SpecPicks estimates derived from the model's architecture, the quant file sizes and memory-bandwidth specifications. Treat them as planning ranges, not measured results.

— Mike Perry · Last verified September 24, 2026

Products mentioned in this article

Live Amazon & eBay pricing, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

Asus ROG Strix Gaming RTX 3090 Review, Thermals, Overclocking & Gaming Benchmarks — Hardware Unboxed on YouTube

Frequently asked questions

How much memory does Mixtral 8x22B need at Q4?
Mixtral 8x22B has about 141B total parameters, so a Q4_K_M GGUF is roughly 80-86 GB before KV cache. That rules out any single consumer GPU for full-VRAM inference. Your options are three or four 24 GB cards, or a mix of GPU VRAM plus 96-128 GB of system RAM, with the expert layers offloaded to the CPU.
Why does expert offload work better for Mixtral than for a dense 70B model?
Only two of the eight experts fire per token, so roughly 39B parameters are active, per Mistral's release post. When the expert tensors sit in system RAM, the CPU reads a fraction of the model per token. A dense 70B model has to read all of its weights every time. That makes MoE offload workable on a 12-24 GB card.
Is Mixtral 8x22B still worth running in 2026?
Newer MoE releases often score higher on reasoning, but Mixtral 8x22B keeps its Apache 2.0 license, 64K context, and mature tooling support in llama.cpp, vLLM, and Ollama. If you need a permissively licensed model with well-understood behavior for production pipelines, it remains a reasonable pick. Otherwise, compare it with current open-weight MoE models first.
How much system RAM do I need alongside the GPU?
For Q4 with expert offload, plan on 128 GB of system RAM if the GPU has 12-16 GB, or 96 GB with a 24-32 GB card. Dual-channel DDR5 bandwidth sets the ceiling on offloaded token speed, so faster rated kits help. Populating four DIMMs on consumer boards often forces lower memory speeds, so check your board's QVL.
Can I run it on a single RTX 3060 12GB at all?
Yes, with llama.cpp's MoE CPU-offload options and enough system RAM, but expect single-digit tokens per second at best. The GPU holds attention and shared layers, and the CPU handles the experts. That's fine for batch summarization or overnight jobs, but too slow for interactive coding. Treat it as a starting point before a 24 GB upgrade.

Sources

— Mike Perry · Last verified 2026-09-24

Parts this article names

Amazon Associate — prices tracked 2026-09-23, may vary.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →