Skip to main content

Used RTX 3090 vs New GPU for Local LLMs: 2026 Reddit Consensus

Nine community threads and published benchmark tables weigh a used 24 GB RTX 3090 against the RTX 5060 Ti, 5070 Ti, 4090, 5090 and dual RTX 3060s.

Used RTX 3090s now run $1,242–$1,395 on eBay, but 24 GB on one card keeps them the community default for 27B–32B local LLMs. When a new card wins.

Used RTX 3090 vs New GPU for Local LLMs: 2026 Reddit Consensus

Hardware at a Glance

Median generation throughput at 7–9B models (Llama 3.1 8B, Qwen 3 8B), Q4 quantization, from community-reported runs SpecPicks tracks. Each row pools runs from different sources, runtimes and models in that class, so the rows are not a matched head-to-head; where the article compares cards on the same rig, its own figures are the like-for-like result. Street price is the second-lowest listing priced within the last 24 hours inside a sane band of MSRP, so no single listing sets it; where too few listings pass that check the row shows launch MSRP instead.

GPUVRAM Llama-3-8B class, Q4Street price Sources
NVIDIA GeForce RTX 5090 32 GB 224.8 tok/s4 runs · 3 sources $1,999MSRP SpecPicks median of 4 runs; sources: llama.cpp GitHub, DatabaseMart, Hardware Corner
NVIDIA GeForce RTX 4090 24 GB 129.2 tok/s8 runs · 7 sources $1,599MSRP SpecPicks median of 8 runs; sources: GPU-Benchmarks-on-LLM-Inference…, Awesome Agents LLM Leaderboard, Database Mart, Hardware Corner +3 more
NVIDIA GeForce RTX 5070 Ti 16 GB 122.8 tok/s6 runs · 6 sources $1,190street, all listings SpecPicks median of 6 runs; sources: Compute Market, ComputingForGeeks, Hardware Corner, KnightLi +2 more
NVIDIA GeForce RTX 3090 24 GB 92 tok/s5 runs · 4 sources $1,900street, all listings SpecPicks median of 5 runs; sources: Hardware Corner, kunalganglani.com LLM Benchmarks, LocalScore.ai, MyAIHardware
NVIDIA GeForce RTX 5060 Ti 16 GB 71 tok/s11 runs · 8 sources $789street, all listings SpecPicks median of 11 runs; sources: LocalScore.ai (Mozilla Builders), RunAIHome, Runyard.dev, ComputingForGeeks +4 more
GeForce RTX 3060 12 GB 12 GB 55 tok/s21 runs · 9 sources $329MSRP SpecPicks median of 21 runs; sources: TYO Lab blog, Ajit Singh / Hardware-Corner, Hardware Corner, llama.cpp GitHub Discussion #10879 +5 more

Which models fit on a RTX 3090?

RTX 3090 carries 24 GB of VRAM. The weights column is the range of real Q4_K_M files on Hugging Face for the models in each size class (about 0.6 GB per billion parameters), and the runtime plus a usable context window wants about 2 GB on top; every tokens-per-second figure is a median over community-reported Q4 runs SpecPicks tracks for this card, with the run count and the source beside it.

Model size Weights at Q4 Fits in 24 GB? Measured Left for context Source
3B (Llama 3.2 3B, Qwen 3 4B)Runs on almost anything with a discrete GPU, and usably on modern integrated graphics. about 2 GB Fitsweights and a usable context window Nothing on file → ~21 GBfor runtime and KV cache —
7-9B (Llama 3.1 8B, Qwen 3 8B)The mainstream local model. An 8 GB card fits it; a 12 GB card fits it with real context. 4.4–5.8 GB Fitsweights and a usable context window 92 tok/s5 runs · 4 sources ~18 GBfor runtime and KV cache MyAIHardware
12-14B (Qwen 3 14B, Phi-4)Where 8 GB stops being enough. This is the band the RTX 3060 12GB exists for. 7.3–9.1 GB Fitsweights and a usable context window 55.4 tok/s6 runs · 4 sources ~14 GBfor runtime and KV cache SmeltCore
20-27B (Gemma 3 27B, Mistral Small)A 27B Q4_K_M file (about 16.7 GB) is more than a 16 GB card holds, so 16 GB means a Q3 quant or partial CPU offload; a 24B loads with a short context. 20 GB holds the class with a usable context window, 24 GB with a long one. about 16.7 GB Fitsweights and a usable context window Nothing on file → ~7 GBfor runtime and KV cache —
30-35B (Qwen 3 32B, QwQ 32B)The step change. A 24 GB card holds this entirely in VRAM; below that it is CPU offload. about 19.9 GB Fitsweights and a usable context window 29.2 tok/s4 runs · 3 sources ~4 GBfor runtime and KV cache Hardware Corner
70B+ (Llama 3.3 70B, Qwen 2.5 72B)One 48 GB card or two 24 GB cards. A 32 GB card runs it only with layers in system RAM. about 42.5 GB Nospills to system RAM — PCIe bandwidth sets the speed — none —

Every RTX 3090 benchmark run, with its source → GPU picks for local LLM — the same table across every card we track How we source these numbers

Prices shown on product cards may vary; check the retailer for the current price.

A used RTX 3090 is still the card most local-LLM threads build around, but it is no longer cheap: per GPU Poet's eBay tracker, the daily average of the three lowest RTX 3090 listings ran from $1,242 to $1,395 in September 2026, up from $691 to $997 in February 2026. What it buys is 24 GB of VRAM on one card (NVIDIA spec page) and, per Hardware Corner's GPU ranking, 52.14 tok/s on Qwen3 14B at 16K context against 32.91 tok/s for an RTX 5060 Ti 16GB. The community answer for October 2026 is a qualified yes. Buy the 3090 if your models need more than 16 GB. Otherwise a new 16 GB Blackwell card is the saner purchase at today's used prices.

This synthesis counts nine public threads from r/LocalLLaMA, r/LocalLLM and Hacker News, posted between September 2025 and July 2026. It sets them against published benchmark tables. For the head-to-head numbers between the two Ampere cards, see the RTX 3060 12GB vs RTX 3090 local LLM comparison. The sibling consensus pages cover the RTX 3060 12GB and running gpt-oss on 12 GB of VRAM.

The consensus: VRAM first, and the 3090 is still the 24 GB default

Six of the nine counted threads are either built on RTX 3090s or recommend one:

  • Dual 3090 is the setup people brag about. In an r/LocalLLaMA thread on a 2x3090 setup, the poster reports roughly 113 tok/s generation and about 4,000 tok/s prompt processing without NVLink after moving from WSL2 to native Ubuntu. They are running Qwen 3.6 27B with 262K context across 48 GB. One reply calls dual 3090 "really the way to go". Another describes the rig as hot and power hungry but working.
  • A single 3090 is enough for serious work. An r/LocalLLM poster runs Qwen3.6-27B at Q4 on one 3090 with 192K context for a Next.js project and calls the results really good.
  • Software keeps squeezing more out of Ampere. The Hacker News thread on 207 tok/s with Qwen3.5-27B on an RTX 3090 covers a speculative-decoding project. Its authors report a 37.78 tok/s autoregressive Q4_K_M baseline on the same card. Commenters dispute how the peak was measured, and the headline 207.6 tok/s is a speculative-decoding peak, not a like-for-like generation speed.
  • Mixed rigs lean on the 3090 for capacity. The HN discussion of an RTX 5080 + RTX 3090 setup pairs a new 16 GB card with a used 24 GB one to run a 27B model at Q8. One commenter reports about 120 tok/s on Qwen3.6-35B-A3B with multi-token prediction on a single 3090.

The published tables back the speed half of that view. Per Hardware Corner (Q4_K_XL, 16K context), the 3090 posts 87.45 tok/s on Qwen3 8B. The RTX 5070 Ti posts 87.54 and the RTX 5060 Ti 16GB posts 51.41. On gpt-oss 20B the figures are 128.51 tok/s for the 3090, 133.05 for the 5070 Ti and 82.42 for the 5060 Ti. LocalScore lists the 3090 at 96.3 tok/s on Llama 3.1 8B Q4_K_M, against 59.0 tok/s for the RTX 5060 Ti. The 3090 matches the 5070 Ti on speed and clearly beats the 5060 Ti. Its one decisive edge over both is 8 GB of extra VRAM.

Hardware Corner's dedicated RTX 3090 page shows what that VRAM buys: 35.1 tok/s on Qwen3 32B at 4K context and 30.3 tok/s at 16K. In the dual-3060 thread, one commenter says a 32B model at Q4_K_M needs about 19 GB. That is more than any 16 GB card can hold without offloading.

The RTX 3090 vs RTX 4090 inference page covers the top of the range. Per Hardware Corner, the 4090 runs Qwen3 14B at 69.14 tok/s at 16K and the 5090 at 102.68 tok/s. The threads treat both as faster, not as better value.

The case for a new card

The new-card side gets a real hearing too:

  • Price has moved against the 3090. In the 207 tok/s HN thread, one commenter says the cheapest working 3090 on eBay came to $1,800 CAD after exchange rate and shipping, against the roughly $700 they once sold for. In the dual-3060 thread, one reply calls 3090s overpriced and suggests an $800–1,000 budget for one. Another reply in the same thread says they have been $550–650 on r/hardwareswap. Asking prices clearly vary by venue and date, and the GPU Poet series is the dated reference here.
  • Power and heat. NVIDIA rates the 3090 at 350 W board power with a 750 W system recommendation (spec page). It rates the RTX 5060 Ti at 180 W total graphics power (NVIDIA). In the 5080 + 3090 thread, a commenter says that kind of build draws 700 W at full load. Another caps each card at 220 W with nvidia-smi and reports losing only about 1 tok/s.
  • Blackwell features. NVIDIA's RTX 5060 family page lists 5th-generation Tensor Cores and FP4 support on the 5060 Ti. The 3090 spec page lists 3rd-generation Tensor Cores. The counted threads mention FP8 and FP4 in passing but do not benchmark them, so this synthesis makes no throughput claim for FP4.
  • 16 GB is enough for mixture-of-experts models. In an r/LocalLLaMA comparison of Qwen3.6 35B-A3B and Gemma 4 26B-A4B, the poster runs both on a 16 GB card at comparable speeds. Mixture-of-experts models are the main reason a 16 GB card no longer feels cramped.
  • New 16 GB stock is not guaranteed. The HN thread on NVIDIA reportedly ending RTX 5070 Ti production (January 2026, with the 5060 Ti 16GB reportedly next) has commenters warning that 16 GB consumer cards may get scarce. That argues for buying new soon, if new is the plan.

Warranty comes up less than one might expect. None of the counted threads puts a figure on it. The practical point holds anyway: a new card comes with recourse, while a used 3090 bought from a stranger usually does not.

Used-buying risks the threads raise

  • Mining history. The dual-3090 rig thread on HN argues both sides. One commenter warns that many used 3090s came off the crypto boom and may be overused. Another reports running seven ex-mining 3090s for two years without trouble. A third says two ex-mining cards bought about three years earlier have had no problems. In the 5080 + 3090 thread, a reply to someone who paid €390 for a second-hand 3090 asks whether it was a mining card.
  • Corroded heatsinks. In the same dual-3090 thread, one commenter notes many eBay 3090 listings showing rusted or corroded heatsinks. That is a visible red flag worth checking in listing photos.
  • Heat, noise and fit. The 2x3090 r/LocalLLaMA thread has a builder whose second 3090 would not fit in their case. The HN dual-3090 thread has commenters who needed a larger case or risers for airflow.
  • Power estimates are contested. One HN commenter estimates 1,400 W for four 3090s at load. Another replies that four 3090s running inference on a large model draw closer to 350 W combined (thread). Actual draw depends heavily on workload and power limits.

None of the counted threads gives measured VRAM junction temperatures or before-and-after thermal-pad results. Treat repadding as a known Ampere maintenance item rather than something these sources quantify.

Dual RTX 3060 12GB vs a single 3090

The r/LocalLLM thread asking whether dual 12 GB 3060s are the right track is the clearest debate. The poster has a $400–600 budget. Most replies favour one 3090. One commenter says the 3060 is half as fast as a 3090 and that 2x12 GB yields less usable VRAM than one 24 GB card because of overhead. Another reports about 48 tok/s on Qwen2.5-14B Q5_K_M on a 3090, against about 19 tok/s on a friend's dual-3060 setup.

The dissent is about budget. One reply reports buying two 3060s for about $400 total from eBay and Facebook Marketplace. That commenter uses the pair for batch work and a 3090 Ti for interactive use, and says two 3060s will not generate tokens like a 3090-class card.

The 207 tok/s HN thread has the strongest pro-3060 voice. That commenter runs Qwen 27B and 35B on two 3060s at 8K–16K context, at 14 tok/s for dense models and 68 tok/s for mixture-of-experts. They say a three-card rig cost less than one 3090, and a reply challenges them to post a full parts list. In the r/LocalLLM budget thread, one reply puts a used RTX 3060 12GB at about $200.

Per Hardware Corner, a single RTX 3060 runs Qwen3 14B at 22.66 tok/s against 52.14 for the 3090. The parts list for the two-card route is in the dual RTX 3060 24GB build guide. What one 3060 can do alone is in the RTX 3060 12GB complete local LLM guide.

Who should buy what

Every thread counted

  1. we really all are going to make it, aren't we? 2x3090 setup. (r/LocalLLaMA, May 2026)
  2. On the right track with dual 12gb 3060's? (r/LocalLLM, July 2026)
  3. Is a 3050/60 a smart choice right now given the rumor that NVIDIA is rereleasing? (r/LocalLLM, March 2026)
  4. Best local ai for coding Nextjs project (r/LocalLLM, May 2026)
  5. Layman's comparison on Qwen3.6 35b-a3b and Gemma4 26b-a4b-it (r/LocalLLaMA, April 2026)
  6. RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8 (Hacker News, June 2026)
  7. We got 207 tok/s with Qwen3.5-27B on an RTX 3090 (Hacker News, April 2026)
  8. 25L Portable NV-linked Dual 3090 LLM Rig (Hacker News, September 2025; the submission links an r/LocalLLaMA build post)
  9. Nvidia Reportedly Ends GeForce RTX 5070 Ti Production, RTX 5060 Ti 16 GB Next (Hacker News, January 2026)

Tally across these nine: six are built on RTX 3090s or recommend one (1, 2, 4, 6, 7, 8). Three make the case for 16 GB or budget cards (3, 5, 9). Two describe 3090 used prices as inflated (2, 7). Two raise mining wear (6, 8).

Citations and sources

  • https://gpupoet.com/gpu/learn/price/september-2026/nvidia-geforce-rtx-3090
  • https://gpupoet.com/gpu/learn/price/february-2026/nvidia-geforce-rtx-3090
  • https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090-3090ti/
  • https://www.hardware-corner.net/gpu-ranking-local-llm/
  • https://www.reddit.com/r/LocalLLaMA/comments/1tcf2dt/we_really_all_are_going_to_make_it_arent_we/
  • https://www.reddit.com/r/LocalLLM/comments/1t9371m/best_local_ai_for_coding_nextjs_project/
  • https://news.ycombinator.com/item?id=47838788
  • https://news.ycombinator.com/item?id=48515454
  • https://www.localscore.ai/accelerator/1
  • https://www.localscore.ai/accelerator/860
  • https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/
  • https://www.reddit.com/r/LocalLLM/comments/1uv4rgo/on_the_right_track_with_dual_12gb_3060s/
  • https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5060-family/
  • https://www.reddit.com/r/LocalLLaMA/comments/1sqxiz0/laymans_comparison_on_qwen36_35ba3b_and_gemma4/
  • https://news.ycombinator.com/item?id=46632955
  • https://news.ycombinator.com/item?id=45300668
  • https://www.reddit.com/r/LocalLLM/comments/1s8cymo/is_a_305060_a_smart_choice_right_now_given_the/

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Amazon & eBay listings, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Watch a review

Asus ROG Strix Gaming RTX 3090 Review, Thermals, Overclocking & Gaming Benchmarks — Hardware Unboxed on YouTube

Frequently asked questions

Is a used RTX 3090 still worth buying for local LLMs in 2026?
Yes, if your models need more than 16 GB of VRAM. Six of the nine threads counted here are built on 3090s or recommend one, for its 24 GB on a single card. Prices have risen, though. GPU Poet's eBay tracker shows the daily three-lowest-listing average at $1,242 to $1,395 in September 2026, up from $691 to $997 in February 2026.
How fast is an RTX 3090 compared with an RTX 5060 Ti 16GB for LLM inference?
Per Hardware Corner's ranking (Q4_K_XL, 16K context), the RTX 3090 runs Qwen3 14B at 52.14 tok/s, against 32.91 tok/s for the RTX 5060 Ti 16GB. On gpt-oss 20B the figures are 128.51 and 82.42 tok/s. LocalScore lists 96.3 tok/s against 59.0 tok/s on Llama 3.1 8B Q4_K_M.
Is the RTX 5070 Ti faster than a used RTX 3090 for local LLMs?
Per Hardware Corner, the two are close. The RTX 5070 Ti runs Qwen3 8B at 87.54 tok/s against 87.45 for the 3090, Qwen3 14B at 57.98 against 52.14, and gpt-oss 20B at 133.05 against 128.51. The 5070 Ti has 16 GB of VRAM to the 3090's 24 GB, so models that need more than 16 GB favour the 3090.
Are two RTX 3060 12GB cards as good as one RTX 3090?
Most replies in the counted r/LocalLLM thread say no. One commenter reports about 48 tok/s on Qwen2.5-14B Q5_K_M on a 3090, against about 19 tok/s on a dual-3060 setup, and notes that 2x12 GB gives less usable VRAM than one 24 GB card. Dual 3060s win only on a strict budget of around $400 for the pair, or for batch work.
What should I check before buying a used RTX 3090?
The counted threads raise three things. Many used 3090s came from crypto mining, and opinion on wear is split. Some eBay listings show rusted or corroded heatsinks. Power and heat are high: NVIDIA rates the card at 350 W, and builders report capping cards at about 220 W with nvidia-smi for only about 1 tok/s lost.

— Mike Perry · Updated 2026-10-08

Parts this article names

Amazon Associate — prices tracked 2026-10-08, may vary.