Skip to main content

gpt-oss on a 12GB GPU: 2026 Community Consensus (RTX 3060)

What 12 public threads report for gpt-oss-20b fully in VRAM and gpt-oss-120b with expert offload to system RAM

gpt-oss-20b runs at 53-85 tok/s on an RTX 3060 12GB; gpt-oss-120b needs --n-cpu-moe and 64GB+ RAM. What 12 community threads report.

Quick answer

Yes. gpt-oss-20b runs on an RTX 3060 12GB at 75 tok/s when the whole model and a 2,048-token context fit in VRAM, and at 64 tok/s with a 16K context once two layers' experts move to the CPU (-ncmoe 2), per an RTX 3060 owner's llama-bench and llama-server results in llama.cpp discussion #15396. gpt-oss-120b also runs on a 12GB card, but only with its experts offloaded to 64 GB or more of system RAM. The closest published 12GB result is 18-22 tok/s on an RTX 3080 Ti 12GB with 128 GB of DDR4-3600 and --n-cpu-moe 36, per David Crook's write-up. The 20b model is the easy part. Whether the 120b model is usable depends on your system RAM and how long your prompts are.

As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices may vary.

This synthesis counts 12 public discussion threads: a llama.cpp GitHub discussion, three Ollama GitHub issues, six Hacker News threads, one r/LocalLLaMA thread and one Habr comment thread. Each is linked in "Every thread counted" below. Six of the 12 involve an RTX 3060 directly. For the two head-to-head pages behind this consensus, see gpt-oss 20B on an RTX 3060 12GB vs a Ryzen 5 5600G CPU-only and gpt-oss 120B with RTX 3060 offload vs a Ryzen 7 5800X CPU-only. The wider card-level picture is in the RTX 3060 12GB complete local LLM guide and the sibling RTX 3060 12GB local LLM Reddit consensus.

gpt-oss on a 12GB GPU: 2026 Community Consensus (RTX 3060)

Hardware at a Glance

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead. Rows marked for comparison are not covered by this article — they are the nearest cards by VRAM, included so the throughput column has something to be read against.

GPUVRAM Llama-3-8B class, Q4Street price Sources
GeForce RTX 3060 12 GB 12 GB 55 tok/s21 runs · 9 sources $329MSRP SpecPicks median of 21 runs; sources: TYO Lab blog, Ajit Singh / Hardware-Corner, Hardware Corner, llama.cpp GitHub Discussion #10879 +5 more
GeForce RTX 4070 SUPERfor comparison 12 GB 60.6 tok/s10 runs · 6 sources $969street, all listings SpecPicks median of 10 runs; sources: llmrun.dev, Hardware Corner, LocalScore.ai, llama.cpp GitHub Discussions +2 more
Arc B580for comparison 12 GB 41 tok/s11 runs · 9 sources $249MSRP SpecPicks median of 11 runs; sources: Compute Market, llama.cpp GitHub Discussions, dev.to, InsiderLLM +5 more

gpt-oss-20b on 12GB: the consensus is "yes, fast"

The MXFP4 GGUF of gpt-oss-20b is 11.27 GiB, per the llama.cpp gpt-oss guide. That leaves very little of a 12GB card for the KV cache. The RTX 3060 owner in that discussion reports 12,288 MB of VRAM with 378 MB reserved by the Linux driver, per the same thread. In practice the 20b model fits completely only at small contexts. Larger contexts push one to three layers' experts into system RAM.

Two threads put a generation speed on gpt-oss-20b on an RTX 3060, and both land in the same range:

  • llama.cpp discussion #15396. At 16K context the 3060 owner reports 64 tok/s on the MXFP4 file with -ncmoe 2, and 67 tok/s on Unsloth's UD-Q4_K_XL with only the later up-projection experts on the CPU (-ot "\.([2-9][0-9])\.ffn_up_exps.=CPU"). At 32K context the figures are 56 tok/s on MXFP4 (-ncmoe 3, which the owner says will likely run out of memory before the context fills) and 53 tok/s on UD-Q4_K_XL, per the RTX 3060 benchmark comment. With nothing offloaded and a 2,048-token window, the same card reaches 75 tok/s.
  • Hacker News, "The $2k Laptop That Replaced My $200/Month AI Subscription". One commenter reports gpt-oss-20b running at 85 t/s "with a smallish context" on a dedicated 12GB 3060, and describes the MXFP4 file as 12.1GB, per the HN thread.

Prompt processing is not the bottleneck for the 20b model on this card. The same llama-bench run shows 2,229.95 tok/s at a 2,048-token prompt and 1,960.34 tok/s at 16,384 tokens with up-projection experts on the CPU, per llama.cpp discussion #15396. A 12GB RTX 5070 posted 62.21 tok/s at 32K context with --n-cpu-moe 2 in the same thread. The two owners agreed in the thread that once anything is offloaded, the CPU and system RAM set the pace rather than the GPU. With nothing offloaded the 5070 reached 128 t/s against the 3060's 75 t/s, per the follow-up exchange.

The community's verdict is that gpt-oss-20b is the default gpt-oss model for a 12GB card. Keep the context at 16K or below if you want generation to stay above 60 tok/s.

gpt-oss-120b with expert offload: what it takes

gpt-oss-120b in MXFP4 is 59.02 GiB, per the llama.cpp guide's benchmark table. No 12GB card holds it, so every working report uses llama.cpp's MoE offload. Attention, KV cache and routing stay on the GPU, and some or all of the expert weights stay in system RAM.

The flags people actually use:

  • --n-cpu-moe N (short form -ncmoe N) keeps the expert weights of the first N layers on the CPU. The guide's example for an RTX 2060 8GB runs gpt-oss-120b at 32K context with --n-cpu-moe 35, and its 16GB example uses --n-cpu-moe 32, per the llama.cpp gpt-oss guide.
  • --cpu-moe puts every expert layer on the CPU. This is the configuration in the r/LocalLLaMA thread "120B runs awesome on just 8GB VRAM!", which GeekNews summarizes as 5 GB of VRAM in use, 122.66 tok/s of prompt processing and 18.04 tok/s of generation.
  • -ot / --override-tensor with a regex moves specific expert tensors to the CPU. A commenter in the HN discussion of that Reddit thread notes it is "just doing a regex on the layer names", so it also works on other MoE models.

Speeds reported on 8-12GB cards. The original poster of the r/LocalLLaMA thread reports 25.62 tok/s of generation and 134.44 tok/s of prompt processing at 8 GB of VRAM with --n-cpu-moe 36, per GeekNews. The 12GB RTX 3080 Ti build reports 18-22 tok/s while using about 6 GB of VRAM and about 60 GB of system RAM, per Crook. On Habr, an RX 6600 8GB with 64 GB of DDR4-3600 gives 13 t/s, per the Habr comment thread. The 12GB RTX 5070 owner in llama.cpp #15396 got 12 t/s from the 120b model but said there was "not enough memory to run the 120B model reliably", per that thread.

System RAM. Three threads give explicit RAM guidance, and all three put the floor at about 64 GB. The r/LocalLLaMA poster calls 64 GB the minimum and 96 GB ideal, per GeekNews. A Hacker News commenter running the model on CPU says it "uses about 60-65gb of memory", per "Tell HN: I cut Claude API costs". The RTX 3060 owner in llama.cpp #15396 says the 20b results make "a good case to upgrade RAM to 64GB" for the 120b model, per the discussion. For a 3060 owner on 32 GB, gpt-oss-120b means buying RAM first. The full offload math is in the gpt-oss 120B RTX 3060 offload breakdown.

Where the threads disagree

Is 64 GB enough? The r/LocalLLaMA poster says yes, as a minimum. The Habr RX 6600 user ran on 64 GB of DDR4. An i7 owner with 64 GB and an 8 GB card reports the model working, per Ask HN: who uses open LLMs locally. On the other side, the 12GB RTX 5070 owner with 64 GB of DDR5-6000 called it unreliable. The practical reading is that 64 GB loads the model, but a browser, IDE and the OS compete for what is left. The r/LocalLLaMA poster's 96 GB "ideal" exists to leave that headroom.

"Fast" or "slow"? A user with a 5950X, 128 GB of RAM and a 12GB 3060 says token generation is "excellent" but that "when the context grows even a little processing of it is super slow", per the HN thread on the 8GB-VRAM post. In the same thread a commenter quotes the Reddit discussion: speed falls "from 25T/s to 18T/s for very long context". The i7 owner in Ask HN reports the 120b model took "over an hour" to finish a 50-question quiz that ChatGPT finished in 6 minutes, while scoring 47/50 against 46/50, per the Ask HN thread. The threads agree that generation is acceptable and long prompts are painful. They disagree on whether that trade is worth it.

The 3.6 t/s outlier. One RTX 3060 owner on Habr reported 3.6 t/s for gpt-oss-120b "regardless of settings" on 51.2 GB of RAM. Later in the thread, the logs show llama.cpp warning that "no usable GPU" was found. The build, installed through brew, had no CUDA support, per the Habr thread. If your 120b speed is in single digits, check that llama.cpp detects your GPU before blaming the card.

Runtime choice: llama.cpp, Ollama, LM Studio

Three of the 12 threads are Ollama GitHub issues, and all three concern gpt-oss not using the GPU the way users expected:

  • ollama/ollama #11772 asks for MoE expert offload. One commenter with a Ryzen 9 7950X, 64 GB of RAM and an RTX 3090 measured about 8.5 tok/s for gpt-oss:120b in Ollama at 16K context, against 29.40 tok/s in llama.cpp with --n-cpu-moe 24, "nearly 3.5x the performance".
  • ollama/ollama #11676 reports Ollama 0.11.0 on Windows running gpt-oss-20b and gpt-oss-120b entirely on the CPU while an RTX 4070 Ti and an RTX 3060 12GB sat idle.
  • ollama/ollama #12197, filed from a server with two RTX 3060s, reports some gpt-oss requests running on the CPU even with the model loaded on the GPU. One request took 17 minutes.

These issues date from 2025, and Ollama has changed since then, so check the current release before relying on them. The pattern in these threads still holds. The best 8-12GB results for the 120b model (25.62 tok/s in the r/LocalLLaMA thread and 13 t/s on Habr) both came from llama.cpp with explicit offload settings. The one 3060 owner who runs the 120b model through Ollama calls generation "excellent" but prompt processing "super slow", per Hacker News. In the HN thread, one commenter running the 20b model through LM Studio reports 50-60 T/s on a 2021 M1 and mentions threads suggesting gpt-oss-20b "on ollama is slow", per Hacker News. None of the 12 threads benchmarks LM Studio on a 12GB NVIDIA card, so this synthesis does not rank it.

When to upgrade: 24GB, or more RAM

If gpt-oss-20b is your model, a 12GB 3060 is enough and an upgrade buys speed. Hardware Corner measured 160.3 tok/s for gpt-oss-20b on an RTX 3090 Ti at 4K context, per its RTX 3090 Ti page. A Hacker News commenter reports 195 tps for the 20b model on a 24GB RTX 3090, per "The Future of AI Software Development".

For gpt-oss-120b, extra VRAM helps less than you might expect. The same commenter gets 25 tps on the 120b model with a 3090, and the Ollama #11772 commenter got 29.40 tok/s on a 3090 with 64 GB of RAM, per that issue. Those figures are close to the 18-22 tok/s from the 12GB 3080 Ti build. With 32 GB of VRAM, one HN user measured 37 tok/s and found gpt-oss-120b "2 times slower" than Qwen3 32B, per the GPT-OSS vs Qwen3 thread. Even the r/LocalLLaMA poster's 22 GB configuration only raised generation from 25.62 to 30.82 tok/s, per GeekNews. For the 120b model, RAM capacity decides whether it runs at all, and the GPU mostly speeds up prompt processing.

The upgrade paths are compared in RTX 3060 12GB vs RTX 3090 for local LLMs, the used RTX 3090 vs new GPU Reddit consensus and the dual RTX 3060 24GB parts list. If the 3060 also handles media transcoding, see Jellyfin NVENC transcoding plus a local LLM on one RTX 3060 12GB.

Who should do what

  • RTX 3060 12GB with 16-32 GB of RAM: run gpt-oss-20b in llama.cpp at up to 16K context with -ncmoe 2. The threads report 53-85 tok/s depending on context length. Skip the 120b model.
  • RTX 3060 12GB with 64 GB of RAM: the 120b model will load with --n-cpu-moe set high (the 8-16GB configurations above use values from 32 to 36). Expect generation in the low tens of tok/s and slow long prompts. Close other memory-heavy apps.
  • Planning a 120b build: buy 96-128 GB of RAM before buying a bigger GPU. Both 12GB-card builds with 128 GB in these sources (the 3060 owner on HN and the 3080 Ti write-up) ran the model without reporting memory trouble.
  • Using Ollama and seeing single-digit tok/s: confirm the GPU is in use, and try llama-server with explicit offload flags. #11772 measured the difference at about 3.5x on a 3090.
  • Wanting fast gpt-oss-20b at long context: a 24GB card such as the RTX 3090 holds the whole model and the KV cache without offload.

Every thread counted

  1. guide : running gpt-oss with llama.cpp (Discussion #15396) — ggml-org/llama.cpp, GitHub Discussions
  2. GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM — Hacker News
  3. 120B runs awesome on just 8GB VRAM! — r/LocalLLaMA (read via the GeekNews summary and the HN discussion above)
  4. The $2k Laptop That Replaced My $200/Month AI Subscription — Hacker News
  5. use cpu to offload moe weights to reduce the VRAM usage. (#11772) — ollama/ollama, GitHub Issues
  6. Ollama not using NVIDIA GPUs with gpt-oss models (#11676) — ollama/ollama, GitHub Issues
  7. Some requests get processed on CPU, even though model is loaded in GPU (GPT-OSS) (#12197) — ollama/ollama, GitHub Issues
  8. Running GPT-OSS-120B on a 6 GB GPU and speeding it up to 30 t/s: comments — Habr
  9. Ask HN: Who uses open LLMs and coding assistants locally? Share setup and laptop — Hacker News
  10. Tell HN: I cut Claude API costs from $70/month to pennies — Hacker News
  11. The Future of AI Software Development — Hacker News
  12. GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2 (comment) — Hacker News

Citations and sources

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Amazon & eBay listings, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

I'm still mad… but buy it anyway - RTX 3060 Review — Linus Tech Tips on YouTube

Frequently asked questions

Can gpt-oss-20b run on an RTX 3060 12GB?
Yes. An RTX 3060 owner in llama.cpp discussion #15396 reports 75 tok/s with a 2,048-token context fully in VRAM and 64 tok/s at 16K context with two layers' experts on the CPU (-ncmoe 2). A Hacker News commenter reports 85 t/s on a 12GB 3060 with a smallish context.
Can gpt-oss-120b run on a 12GB GPU?
Yes, with MoE expert offload to system RAM in llama.cpp (--n-cpu-moe, --cpu-moe or -ot). The closest published 12GB result is 18-22 tok/s on an RTX 3080 Ti 12GB with 128 GB of DDR4-3600 and --n-cpu-moe 36, using about 6 GB of VRAM and about 60 GB of system RAM.
How much system RAM does gpt-oss-120b need with a 12GB card?
The threads put the floor at about 64 GB. The MXFP4 file is 59.02 GiB, a CPU-only user reports 60-65 GB of memory use, and the r/LocalLLaMA poster calls 64 GB the minimum and 96 GB ideal. A 12GB RTX 5070 owner with 64 GB of DDR5 said it was not enough to run the 120b model reliably.
Is Ollama or llama.cpp better for gpt-oss on a 12GB GPU?
The community reports favor llama.cpp with explicit offload flags. In Ollama issue #11772, an RTX 3090 user measured about 8.5 tok/s for gpt-oss:120b in Ollama against 29.40 tok/s in llama.cpp with --n-cpu-moe 24. Two other 2025 Ollama issues report gpt-oss work landing on the CPU on systems with RTX 3060 cards.
Does a 24GB card like the RTX 3090 make gpt-oss-120b much faster?
Only modestly. RTX 3090 users report 25 tps and 29.40 tok/s for gpt-oss-120b, close to the 18-22 tok/s from a 12GB RTX 3080 Ti build. The bigger gain is for gpt-oss-20b, where a 3090 user reports 195 tps and Hardware Corner measured 160.3 tok/s on an RTX 3090 Ti.

— Mike Perry · Updated 2026-10-08

Parts this article names

Amazon Associate — prices tracked 2026-10-08, may vary.