gpt-oss on a 12GB GPU: 2026 Community Consensus (RTX 3060)
What 12 public threads report for gpt-oss-20b fully in VRAM and gpt-oss-120b with expert offload to system RAM
By Mike Perry, Founder & Editor-in-Chief · Published 2026-10-08 · Updated 2026-10-08 · 10 min read
gpt-oss-20b runs at 53-85 tok/s on an RTX 3060 12GB; gpt-oss-120b needs --n-cpu-moe and 64GB+ RAM. What 12 community threads report.
Quick answer
Yes. gpt-oss-20b runs on an RTX 3060 12GB at 75 tok/s when the whole model and a 2,048-token context fit in VRAM, and at 64 tok/s with a 16K context once two layers' experts move to the CPU (-ncmoe 2), per an RTX 3060 owner's llama-bench and llama-server results in llama.cpp discussion #15396. gpt-oss-120b also runs on a 12GB card, but only with its experts offloaded to 64 GB or more of system RAM. The closest published 12GB result is 18-22 tok/s on an RTX 3080 Ti 12GB with 128 GB of DDR4-3600 and --n-cpu-moe 36, per David Crook's write-up. The 20b model is the easy part. Whether the 120b model is usable depends on your system RAM and how long your prompts are.
As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices may vary.
Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead. Rows marked for comparison are not covered by this article — they are the nearest cards by VRAM, included so the throughput column has something to be read against.
The MXFP4 GGUF of gpt-oss-20b is 11.27 GiB, per the llama.cpp gpt-oss guide. That leaves very little of a 12GB card for the KV cache. The RTX 3060 owner in that discussion reports 12,288 MB of VRAM with 378 MB reserved by the Linux driver, per the same thread. In practice the 20b model fits completely only at small contexts. Larger contexts push one to three layers' experts into system RAM.
Two threads put a generation speed on gpt-oss-20b on an RTX 3060, and both land in the same range:
llama.cpp discussion #15396. At 16K context the 3060 owner reports 64 tok/s on the MXFP4 file with -ncmoe 2, and 67 tok/s on Unsloth's UD-Q4_K_XL with only the later up-projection experts on the CPU (-ot "\.([2-9][0-9])\.ffn_up_exps.=CPU"). At 32K context the figures are 56 tok/s on MXFP4 (-ncmoe 3, which the owner says will likely run out of memory before the context fills) and 53 tok/s on UD-Q4_K_XL, per the RTX 3060 benchmark comment. With nothing offloaded and a 2,048-token window, the same card reaches 75 tok/s.
Hacker News, "The $2k Laptop That Replaced My $200/Month AI Subscription". One commenter reports gpt-oss-20b running at 85 t/s "with a smallish context" on a dedicated 12GB 3060, and describes the MXFP4 file as 12.1GB, per the HN thread.
Prompt processing is not the bottleneck for the 20b model on this card. The same llama-bench run shows 2,229.95 tok/s at a 2,048-token prompt and 1,960.34 tok/s at 16,384 tokens with up-projection experts on the CPU, per llama.cpp discussion #15396. A 12GB RTX 5070 posted 62.21 tok/s at 32K context with --n-cpu-moe 2 in the same thread. The two owners agreed in the thread that once anything is offloaded, the CPU and system RAM set the pace rather than the GPU. With nothing offloaded the 5070 reached 128 t/s against the 3060's 75 t/s, per the follow-up exchange.
The community's verdict is that gpt-oss-20b is the default gpt-oss model for a 12GB card. Keep the context at 16K or below if you want generation to stay above 60 tok/s.
gpt-oss-120b with expert offload: what it takes
gpt-oss-120b in MXFP4 is 59.02 GiB, per the llama.cpp guide's benchmark table. No 12GB card holds it, so every working report uses llama.cpp's MoE offload. Attention, KV cache and routing stay on the GPU, and some or all of the expert weights stay in system RAM.
The flags people actually use:
--n-cpu-moe N (short form -ncmoe N) keeps the expert weights of the first N layers on the CPU. The guide's example for an RTX 2060 8GB runs gpt-oss-120b at 32K context with --n-cpu-moe 35, and its 16GB example uses --n-cpu-moe 32, per the llama.cpp gpt-oss guide.
--cpu-moe puts every expert layer on the CPU. This is the configuration in the r/LocalLLaMA thread "120B runs awesome on just 8GB VRAM!", which GeekNews summarizes as 5 GB of VRAM in use, 122.66 tok/s of prompt processing and 18.04 tok/s of generation.
-ot / --override-tensor with a regex moves specific expert tensors to the CPU. A commenter in the HN discussion of that Reddit thread notes it is "just doing a regex on the layer names", so it also works on other MoE models.
Speeds reported on 8-12GB cards. The original poster of the r/LocalLLaMA thread reports 25.62 tok/s of generation and 134.44 tok/s of prompt processing at 8 GB of VRAM with --n-cpu-moe 36, per GeekNews. The 12GB RTX 3080 Ti build reports 18-22 tok/s while using about 6 GB of VRAM and about 60 GB of system RAM, per Crook. On Habr, an RX 6600 8GB with 64 GB of DDR4-3600 gives 13 t/s, per the Habr comment thread. The 12GB RTX 5070 owner in llama.cpp #15396 got 12 t/s from the 120b model but said there was "not enough memory to run the 120B model reliably", per that thread.
System RAM. Three threads give explicit RAM guidance, and all three put the floor at about 64 GB. The r/LocalLLaMA poster calls 64 GB the minimum and 96 GB ideal, per GeekNews. A Hacker News commenter running the model on CPU says it "uses about 60-65gb of memory", per "Tell HN: I cut Claude API costs". The RTX 3060 owner in llama.cpp #15396 says the 20b results make "a good case to upgrade RAM to 64GB" for the 120b model, per the discussion. For a 3060 owner on 32 GB, gpt-oss-120b means buying RAM first. The full offload math is in the gpt-oss 120B RTX 3060 offload breakdown.
Where the threads disagree
Is 64 GB enough? The r/LocalLLaMA poster says yes, as a minimum. The Habr RX 6600 user ran on 64 GB of DDR4. An i7 owner with 64 GB and an 8 GB card reports the model working, per Ask HN: who uses open LLMs locally. On the other side, the 12GB RTX 5070 owner with 64 GB of DDR5-6000 called it unreliable. The practical reading is that 64 GB loads the model, but a browser, IDE and the OS compete for what is left. The r/LocalLLaMA poster's 96 GB "ideal" exists to leave that headroom.
"Fast" or "slow"? A user with a 5950X, 128 GB of RAM and a 12GB 3060 says token generation is "excellent" but that "when the context grows even a little processing of it is super slow", per the HN thread on the 8GB-VRAM post. In the same thread a commenter quotes the Reddit discussion: speed falls "from 25T/s to 18T/s for very long context". The i7 owner in Ask HN reports the 120b model took "over an hour" to finish a 50-question quiz that ChatGPT finished in 6 minutes, while scoring 47/50 against 46/50, per the Ask HN thread. The threads agree that generation is acceptable and long prompts are painful. They disagree on whether that trade is worth it.
The 3.6 t/s outlier. One RTX 3060 owner on Habr reported 3.6 t/s for gpt-oss-120b "regardless of settings" on 51.2 GB of RAM. Later in the thread, the logs show llama.cpp warning that "no usable GPU" was found. The build, installed through brew, had no CUDA support, per the Habr thread. If your 120b speed is in single digits, check that llama.cpp detects your GPU before blaming the card.
Runtime choice: llama.cpp, Ollama, LM Studio
Three of the 12 threads are Ollama GitHub issues, and all three concern gpt-oss not using the GPU the way users expected:
ollama/ollama #11772 asks for MoE expert offload. One commenter with a Ryzen 9 7950X, 64 GB of RAM and an RTX 3090 measured about 8.5 tok/s for gpt-oss:120b in Ollama at 16K context, against 29.40 tok/s in llama.cpp with --n-cpu-moe 24, "nearly 3.5x the performance".
ollama/ollama #11676 reports Ollama 0.11.0 on Windows running gpt-oss-20b and gpt-oss-120b entirely on the CPU while an RTX 4070 Ti and an RTX 3060 12GB sat idle.
ollama/ollama #12197, filed from a server with two RTX 3060s, reports some gpt-oss requests running on the CPU even with the model loaded on the GPU. One request took 17 minutes.
These issues date from 2025, and Ollama has changed since then, so check the current release before relying on them. The pattern in these threads still holds. The best 8-12GB results for the 120b model (25.62 tok/s in the r/LocalLLaMA thread and 13 t/s on Habr) both came from llama.cpp with explicit offload settings. The one 3060 owner who runs the 120b model through Ollama calls generation "excellent" but prompt processing "super slow", per Hacker News. In the HN thread, one commenter running the 20b model through LM Studio reports 50-60 T/s on a 2021 M1 and mentions threads suggesting gpt-oss-20b "on ollama is slow", per Hacker News. None of the 12 threads benchmarks LM Studio on a 12GB NVIDIA card, so this synthesis does not rank it.
When to upgrade: 24GB, or more RAM
If gpt-oss-20b is your model, a 12GB 3060 is enough and an upgrade buys speed. Hardware Corner measured 160.3 tok/s for gpt-oss-20b on an RTX 3090 Ti at 4K context, per its RTX 3090 Ti page. A Hacker News commenter reports 195 tps for the 20b model on a 24GB RTX 3090, per "The Future of AI Software Development".
For gpt-oss-120b, extra VRAM helps less than you might expect. The same commenter gets 25 tps on the 120b model with a 3090, and the Ollama #11772 commenter got 29.40 tok/s on a 3090 with 64 GB of RAM, per that issue. Those figures are close to the 18-22 tok/s from the 12GB 3080 Ti build. With 32 GB of VRAM, one HN user measured 37 tok/s and found gpt-oss-120b "2 times slower" than Qwen3 32B, per the GPT-OSS vs Qwen3 thread. Even the r/LocalLLaMA poster's 22 GB configuration only raised generation from 25.62 to 30.82 tok/s, per GeekNews. For the 120b model, RAM capacity decides whether it runs at all, and the GPU mostly speeds up prompt processing.
RTX 3060 12GB with 16-32 GB of RAM: run gpt-oss-20b in llama.cpp at up to 16K context with -ncmoe 2. The threads report 53-85 tok/s depending on context length. Skip the 120b model.
RTX 3060 12GB with 64 GB of RAM: the 120b model will load with --n-cpu-moe set high (the 8-16GB configurations above use values from 32 to 36). Expect generation in the low tens of tok/s and slow long prompts. Close other memory-heavy apps.
Planning a 120b build: buy 96-128 GB of RAM before buying a bigger GPU. Both 12GB-card builds with 128 GB in these sources (the 3060 owner on HN and the 3080 Ti write-up) ran the model without reporting memory trouble.
Using Ollama and seeing single-digit tok/s: confirm the GPU is in use, and try llama-server with explicit offload flags. #11772 measured the difference at about 3.5x on a 3090.
Wanting fast gpt-oss-20b at long context: a 24GB card such as the RTX 3090 holds the whole model and the KV cache without offload.
As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.
📹 Watch a review
I'm still mad… but buy it anyway - RTX 3060 Review — Linus Tech Tips on YouTube
Frequently asked questions
Can gpt-oss-20b run on an RTX 3060 12GB?
Yes. An RTX 3060 owner in llama.cpp discussion #15396 reports 75 tok/s with a 2,048-token context fully in VRAM and 64 tok/s at 16K context with two layers' experts on the CPU (-ncmoe 2). A Hacker News commenter reports 85 t/s on a 12GB 3060 with a smallish context.
Can gpt-oss-120b run on a 12GB GPU?
Yes, with MoE expert offload to system RAM in llama.cpp (--n-cpu-moe, --cpu-moe or -ot). The closest published 12GB result is 18-22 tok/s on an RTX 3080 Ti 12GB with 128 GB of DDR4-3600 and --n-cpu-moe 36, using about 6 GB of VRAM and about 60 GB of system RAM.
How much system RAM does gpt-oss-120b need with a 12GB card?
The threads put the floor at about 64 GB. The MXFP4 file is 59.02 GiB, a CPU-only user reports 60-65 GB of memory use, and the r/LocalLLaMA poster calls 64 GB the minimum and 96 GB ideal. A 12GB RTX 5070 owner with 64 GB of DDR5 said it was not enough to run the 120b model reliably.
Is Ollama or llama.cpp better for gpt-oss on a 12GB GPU?
The community reports favor llama.cpp with explicit offload flags. In Ollama issue #11772, an RTX 3090 user measured about 8.5 tok/s for gpt-oss:120b in Ollama against 29.40 tok/s in llama.cpp with --n-cpu-moe 24. Two other 2025 Ollama issues report gpt-oss work landing on the CPU on systems with RTX 3060 cards.
Does a 24GB card like the RTX 3090 make gpt-oss-120b much faster?
Only modestly. RTX 3090 users report 25 tps and 29.40 tok/s for gpt-oss-120b, close to the 18-22 tok/s from a 12GB RTX 3080 Ti build. The bigger gain is for gpt-oss-20b, where a 3090 user reports 195 tps and Hardware Corner measured 160.3 tok/s on an RTX 3090 Ti.