As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices shown are catalog prices at the time of writing and may vary. See our review methodology.
Quick answer: The best GPU for gpt-oss 20B is a 16 GB card, and the PNY GeForce RTX 4060 Ti 16GB Verto is the best overall pick. OpenAI's model card says gpt-oss-20b runs "within 16GB of memory" (Hugging Face). Hardware Corner measures 63.2 tok/s on the 4060 Ti 16GB at 4K context, with the whole model in VRAM up to 128K tokens. A 12 GB RTX 3060 also works if you offload a few expert layers.
Why gpt-oss 20B is easier to run than its size suggests
gpt-oss-20b is OpenAI's smaller open-weight model, released under the Apache 2.0 license. It has 20.91B total parameters, but it is a mixture-of-experts (MoE) model: each of its 24 layers holds 32 experts, and only the top 4 run for any given token. That means only 3.61B parameters are active per token, according to the gpt-oss model card on arXiv. OpenAI also ships the MoE weights, which account for more than 90% of the parameters, pre-quantized to MXFP4 at 4.25 bits per parameter. The result is a 12.11 GB GGUF file (ggml-org/gpt-oss-20b-GGUF), or a 14 GB download on Ollama.
Two things follow for your GPU choice. First, 16 GB is the clean fit. The file, the runtime buffers and a long KV cache all fit on one 16 GB card, with no tuning. Second, 12 GB works better than you'd expect. Because only a small slice of the experts fires per token, llama.cpp can park some expert layers in system RAM with a modest speed penalty. That trick would cripple a dense 20B model. It even makes a 4 GB card usable as a stopgap, with most experts on the CPU. What doesn't help much is going bigger than 16 GB just for this model. A 24 GB card is faster because of its bandwidth, not its capacity.
The winner for most buyers is the RTX 4060 Ti 16GB: full-VRAM fit at every context length, 165 W, and a mid-range price. Here's how the rest of the field stacks up.
Step 0 — 12 GB, 16 GB or 24 GB?
| VRAM tier | Fits fully? | Offload needed | Expected experience |
|---|---|---|---|
| 4 GB (GTX 1050 Ti) | No | Nearly all experts to CPU (--cpu-moe) | Slow but usable; limited by system RAM bandwidth |
| 8 GB | No | Many expert layers (--n-cpu-moe in the teens) | Usable chat speed with a fast CPU and RAM |
| 12 GB (RTX 3060, Arc B580) | Almost | 2–3 expert layers | ~56–64 tok/s on a 3060 at 16–32K context |
| 16 GB (RTX 4060 Ti 16GB) | Yes, up to 128K context | None | 63.2 tok/s at 4K, 31.1 tok/s at 128K |
| 24 GB (RTX 4090) | Yes, with room to spare | None | 190.6 tok/s at 4K |
Sources: Hardware Corner's RTX 4060 Ti 16GB and RTX 4090 pages; RTX 3060 figures from a user report in llama.cpp discussion #15396. The 8 GB row follows the offload guidance in the same discussion.
gpt-oss 20B GPU picks at a glance
| Pick | Best For | Key Spec | Price Range | Verdict |
|---|---|---|---|---|
| PNY RTX 4060 Ti 16GB Verto | Best overall | 16 GB, 288 GB/s, 165 W | ~$500–$550 | Whole model in VRAM at any context |
| MSI RTX 3060 12GB | Best value | 12 GB, 360 GB/s, 170 W | ~$290 used, more new | 2-layer offload, ~64 tok/s |
| ASRock Arc B580 Steel Legend 12GB | Intel/Vulkan builders | 12 GB, 456 GB/s, 190 W | ~$310 | Cheap and fast on paper; less mature software |
| MSI RTX 4090 24GB | Best performance | 24 GB, 1,008 GB/s, 450 W | $2,500+ | ~3× the 4060 Ti's speed |
| EVGA GTX 1050 Ti FTW 4GB | Budget stopgap | 4 GB, 112 GB/s | Buy used only | Experts on CPU; end-of-life drivers |
🏆 Best Overall: PNY GeForce RTX 4060 Ti 16GB Verto
Spec chips: 16 GB GDDR6 · 128-bit · 288 GB/s · 165 W TGP · 550 W recommended PSU · $499 launch MSRP
The PNY RTX 4060 Ti 16GB Verto is the cheapest current NVIDIA card that holds all of gpt-oss-20b in VRAM. Hardware Corner describes it "running gpt-oss 20B (MXFP4) with up to 128K context fully in VRAM" and measures 63.2 / 57.8 / 51.5 / 41.1 / 31.1 tok/s generation at 4K / 16K / 32K / 64K / 128K context (Hardware Corner). 63 tok/s is several times faster than you can read, and the model keeps working at the full 128K window instead of falling off a cliff.
NVIDIA specifies 16 GB of GDDR6 on a 128-bit bus, 165 W, and a 550 W system PSU (NVIDIA RTX 4060 family). The 16 GB variant launched on July 18, 2023 at $499 (Wikipedia).
- Pros: no offload tuning at all; 128K context fits; low power; two-slot card fits almost any case.
- Cons: the 128-bit bus (288 GB/s) is slower than a 3060's 360 GB/s, so the win comes from fitting the model, not from raw bandwidth. The 8 GB 4060 Ti is sold under the same name, so check that the listing says 16GB.
💰 Best Value: MSI Gaming GeForce RTX 3060 12GB
Spec chips: 12 GB GDDR6 · 192-bit · 360 GB/s · 170 W · $329 launch MSRP
The MSI Gaming RTX 3060 12GB can't quite hold the whole model, but it barely needs to offload. In llama.cpp discussion #15396, a user running the MXFP4 GGUF on a 3060 with flash attention, a 16K context and -ncmoe 2 (the first two layers' experts on the CPU) reported "64 tok/sec initial generation rate". At 32K with -ncmoe 3 they reported 56 tok/s, with a warning that the setup "will likely OOM". That's on a Ryzen 7 5700X host.
In other words, a 3060 roughly matches the 4060 Ti at 16K context (different test setups, so treat it as ballpark), because its wider bus makes up for the small offload. The trade-off is headroom. Above ~32K context you offload more layers and speed falls. Used 3060s averaged $293 over 30 days on eBay sold listings as of September 2026 (getpcparts). New stock costs more, so check the live price.
- Pros: cheapest route to ~60 tok/s; mature CUDA support; huge used supply.
- Cons: needs
--n-cpu-moetuning; long contexts push more of the model to RAM; pair it with 32 GB of system RAM.
🎯 Best for Intel/Vulkan builders: ASRock Intel Arc B580 Steel Legend 12GB
Spec chips: 12 GB GDDR6 · 192-bit · 456 GB/s · 190 W · $249 launch MSRP
On paper the ASRock Arc B580 Steel Legend 12GB is the strongest 12 GB card here. Its 456 GB/s beats the 3060's 360 GB/s, and it launched at $249 (Wikipedia: Intel Arc). In practice, Intel's software stack still trails CUDA. On the llama.cpp Vulkan scoreboard (Llama 2 7B Q4_0), the B580 generates at 70.14 tok/s vs the RTX 3060's 75.94, and processes prompts at 620.94 vs 1,815.70 tok/s (llama.cpp discussion #10879).
No published gpt-oss-20b measurement for the B580 turned up in this research, so treat it like the 3060: same 12 GB capacity, same 2–3-layer expert offload, and roughly similar generation speed based on the proxy numbers above.
- Pros: lowest price per GB of bandwidth; runs llama.cpp via SYCL or Vulkan; a good fit if you already run Linux with Intel tooling.
- Cons: slower prompt processing; new model architectures often reach CUDA first; fewer ready-made containers.
⚡ Best Performance: MSI Gaming GeForce RTX 4090 24GB
Spec chips: 24 GB GDDR6X · 384-bit · 1,008 GB/s · 450 W · 850 W recommended PSU
If speed matters (agent loops that make hundreds of calls, batch summarization, several users on one box), the MSI RTX 4090 SUPRIM Liquid X 24G is in a different class. Hardware Corner measures 190.6 tok/s at 4K and 70.3 tok/s at 128K (Hardware Corner). The #15396 guide author reported tg128 of 221.95 tok/s and pp2048 of 8,022 tok/s on a 4090, and a later 2026 run in the same thread reached 254.63 tok/s generation.
That's about 3× the 4060 Ti at every context length, for roughly 5× the price and 2.7× the power (450 W per NVIDIA). The 24 GB also leaves room to run a second model beside gpt-oss-20b, or to step up to 27B–32B dense models later.
- Pros: the fastest pick in this guide (only an RTX 5090, at 298.2 tok/s at 4K per Hardware Corner, is quicker); headroom for other models.
- Cons: used prices near $2,500 in September 2026; 850 W PSU; far more speed than single-user chat needs.
🧪 Budget Pick: EVGA GeForce GTX 1050 Ti FTW 4GB
Spec chips: 4 GB GDDR5 · 128-bit · 112 GB/s · $139 launch MSRP (2016)
The EVGA GTX 1050 Ti FTW is here for readers who already own one, or who can get one very cheaply. It is not a card to buy new for AI. With llama.cpp's --cpu-moe flag, which keeps "all Mixture of Experts (MoE) weights in the CPU" per the llama.cpp server README, the GPU holds the attention layers and KV cache while system RAM holds the experts. Generation speed is then set mainly by your RAM bandwidth.
No public gpt-oss-20b measurement exists for a 1050 Ti. As a reference point, a CPU-only Ryzen 9 7950X runs gpt-oss-20b at about 21.7 tok/s at zero context and about 10.4 tok/s at ~24.5K context (ik_llama.cpp discussion #758). A 1050 Ti mainly speeds up prompt processing on top of a CPU-bound setup. On an older DDR4 desktop, expect single-digit to low-teens tok/s (SpecPicks estimate).
Two caveats. First, Pascal is at the end of its software life. NVIDIA's CUDA 13.0 notes say offline compilation and library support for Maxwell, Pascal and Volta "have been removed" (CUDA 13.0 release notes), and the R580 branch is the last Linux driver for those GPUs (Phoronix). Pin a CUDA 12.x build. Second, this FTW model takes a 6-pin power connector, unlike the slot-powered reference card. Catalog listings for new 1050 Ti stock are often priced far above their value, so buy used.
- When it's right: you already own it and want to try gpt-oss-20b today.
- When it isn't: any new purchase. A used RTX 3060 costs little more and is several times faster.
Also considered: ZOTAC Gaming GeForce RTX 3060 Twin Edge OC 12GB
The ZOTAC RTX 3060 Twin Edge OC 12GB uses the same GA106 GPU, 12 GB and 192-bit bus as the MSI pick, in a shorter dual-fan card that suits small cases. Expect the same gpt-oss-20b behavior. Choose between them on price and size, and avoid the 8 GB RTX 3060 variant, which has a narrower 128-bit bus and needs far more expert offload than the 12 GB card.
What to look for in a GPU for gpt-oss 20B
VRAM tier
16 GB avoids offload entirely. 12 GB needs a couple of expert layers on the CPU. Below 12 GB you're running a CPU-bound model with GPU assistance. The model's native context is 131,072 tokens using YaRN (arXiv model card), and on a 16 GB card you can use all of it.
Memory bandwidth
Once the model fits, generation speed tracks bandwidth: 288 GB/s (4060 Ti 16GB), 360 GB/s (3060), 456 GB/s (B580), 1,008 GB/s (4090). The 4090's ~3× advantage over the 4060 Ti in Hardware Corner's numbers matches its ~3.5× bandwidth advantage closely.
MXFP4 and quant support
The official GGUF is MXFP4, which current llama.cpp, Ollama and LM Studio builds support. Community re-quants (for example Unsloth's Q4_K_XL) exist, and the 3060 user in #15396 reported 67 tok/s with one at 16K. You gain little in size, because the experts are already about 4 bits.
Runtime support
llama.cpp exposes --cpu-moe and --n-cpu-moe N for expert offload. The #15396 guide shows an 8 GB RTX 2060 running full context with --n-cpu-moe 22. Ollama handles offload automatically but gives you less control. vLLM expects the whole model in VRAM, so it is a 16 GB+ option.
PSU and power
165 W (4060 Ti 16GB), 170 W (3060), 190 W (B580), 450 W (4090). The first three run on a quality 550–650 W unit; the 4090 needs 850 W.
System RAM for offload
Any card below 16 GB puts experts in system RAM. 32 GB of dual-channel memory is the practical floor, and faster RAM directly speeds up offloaded layers.
FAQ
How much VRAM does gpt-oss 20B need? OpenAI says gpt-oss-20b is designed to run within 16 GB of memory using its native MXFP4 quantization. A 16 GB card holds the whole model plus a moderate context window. A 12 GB card can run it by offloading a few MoE expert layers to system RAM, which costs some speed but keeps the experience usable for single-user chat.
Is the RTX 3060 12GB good enough for gpt-oss 20B? For single-user chat and coding help, generally yes. Only about 3.6B parameters are active per token, so partial offload hurts less than it would on a dense 20B model. You'll see lower tokens per second than on a 16 GB card at long context, and very long contexts will push more of the model into system RAM. Pair it with 32 GB of system memory.
Can I run gpt-oss 20B on a 4 GB card like the GTX 1050 Ti? Only with most of the model in system RAM. llama.cpp can keep the attention layers on the GPU and push the MoE experts to the CPU, which works because so few parameters are active per token. Expect slow but usable output on a fast DDR4 system. Treat it as a stopgap before upgrading to a 12 GB or 16 GB card.
Do Intel Arc cards run gpt-oss 20B? Yes, through llama.cpp's SYCL or Vulkan back-ends and Intel's own inference tooling. The Arc B580's 12 GB puts it in the same offload position as the RTX 3060. The trade-off is software maturity: CUDA builds usually get new model support first. Choose Arc for price and bandwidth, and budget extra time for setup.
How much system RAM should I pair with the GPU? 32 GB is the sensible floor for any build that offloads gpt-oss experts to the CPU, and 64 GB gives room for longer contexts and other models alongside it. Dual-channel memory matters because offloaded layers are limited by RAM bandwidth. On a 16 GB or 24 GB card that holds the full model, 32 GB is plenty.
Citations and sources
- OpenAI, gpt-oss-20b model card on Hugging Face — 21B/3.6B parameters, MXFP4, "within 16GB of memory", Apache 2.0. Accessed September 24, 2026.
- OpenAI, gpt-oss-120b & gpt-oss-20b Model Card (arXiv 2508.10925) — 20.91B total / 3.61B active, 32 experts top-4, 4.25 bits per parameter, 131,072-token context. Accessed September 24, 2026.
- ggml-org, gpt-oss-20b GGUF — 12.11 GB MXFP4 file; Ollama, gpt-oss library page — 14 GB. Accessed September 24, 2026.
- Hardware Corner, RTX 4060 Ti 16GB and RTX 4090 LLM benchmarks — gpt-oss 20B generation by context length. Accessed September 24, 2026.
- ggml-org, llama.cpp gpt-oss guide, discussion #15396 — RTX 3060 and RTX 4090 user results,
--n-cpu-moeguidance. Accessed September 24, 2026. - ggml-org, llama.cpp Vulkan scoreboard, discussion #10879 — Arc B580 vs RTX 3060 proxy results. Accessed September 24, 2026.
- ggml-org, llama.cpp server README —
--cpu-moe/--n-cpu-moeflags. Accessed September 24, 2026. - ikawrakow, ik_llama.cpp discussion #758 — CPU-only gpt-oss-20B on a Ryzen 9 7950X. Accessed September 24, 2026.
- NVIDIA, RTX 4060 family and RTX 4090 specifications; Wikipedia, GeForce RTX 40 series and Intel Arc — specs and launch MSRPs. Accessed September 24, 2026.
- NVIDIA, CUDA 13.0 release notes, and Phoronix, R580 is the last driver for Pascal. Accessed September 24, 2026.
- getpcparts, RTX 3060 used market prices — eBay sold-listing averages as of September 19, 2026. Accessed September 24, 2026.
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported. VRAM-tier guidance and the GTX 1050 Ti speed range are SpecPicks estimates from the cited specifications and measurements.
Related guides
- Best 16GB GPU for Local LLMs in 2026
- Best 12GB GPU for Local LLMs in 2026
- gpt-oss 20B: RTX 3060 12GB vs Ryzen 5 5600G CPU
- Best GPU for Qwen3 8B in 2026
- RTX 4090 vs RTX 5090 for Local LLM Inference
— Mike Perry
