Prices shown on product cards may vary; check the retailer for the current price.
A used RTX 3090 is still the card most local-LLM threads build around, but it is no longer cheap: per GPU Poet's eBay tracker, the daily average of the three lowest RTX 3090 listings ran from $1,242 to $1,395 in September 2026, up from $691 to $997 in February 2026. What it buys is 24 GB of VRAM on one card (NVIDIA spec page) and, per Hardware Corner's GPU ranking, 52.14 tok/s on Qwen3 14B at 16K context against 32.91 tok/s for an RTX 5060 Ti 16GB. The community answer for October 2026 is a qualified yes. Buy the 3090 if your models need more than 16 GB. Otherwise a new 16 GB Blackwell card is the saner purchase at today's used prices.
This synthesis counts nine public threads from r/LocalLLaMA, r/LocalLLM and Hacker News, posted between September 2025 and July 2026. It sets them against published benchmark tables. For the head-to-head numbers between the two Ampere cards, see the RTX 3060 12GB vs RTX 3090 local LLM comparison. The sibling consensus pages cover the RTX 3060 12GB and running gpt-oss on 12 GB of VRAM.
The consensus: VRAM first, and the 3090 is still the 24 GB default
Six of the nine counted threads are either built on RTX 3090s or recommend one:
- Dual 3090 is the setup people brag about. In an r/LocalLLaMA thread on a 2x3090 setup, the poster reports roughly 113 tok/s generation and about 4,000 tok/s prompt processing without NVLink after moving from WSL2 to native Ubuntu. They are running Qwen 3.6 27B with 262K context across 48 GB. One reply calls dual 3090 "really the way to go". Another describes the rig as hot and power hungry but working.
- A single 3090 is enough for serious work. An r/LocalLLM poster runs Qwen3.6-27B at Q4 on one 3090 with 192K context for a Next.js project and calls the results really good.
- Software keeps squeezing more out of Ampere. The Hacker News thread on 207 tok/s with Qwen3.5-27B on an RTX 3090 covers a speculative-decoding project. Its authors report a 37.78 tok/s autoregressive Q4_K_M baseline on the same card. Commenters dispute how the peak was measured, and the headline 207.6 tok/s is a speculative-decoding peak, not a like-for-like generation speed.
- Mixed rigs lean on the 3090 for capacity. The HN discussion of an RTX 5080 + RTX 3090 setup pairs a new 16 GB card with a used 24 GB one to run a 27B model at Q8. One commenter reports about 120 tok/s on Qwen3.6-35B-A3B with multi-token prediction on a single 3090.
The published tables back the speed half of that view. Per Hardware Corner (Q4_K_XL, 16K context), the 3090 posts 87.45 tok/s on Qwen3 8B. The RTX 5070 Ti posts 87.54 and the RTX 5060 Ti 16GB posts 51.41. On gpt-oss 20B the figures are 128.51 tok/s for the 3090, 133.05 for the 5070 Ti and 82.42 for the 5060 Ti. LocalScore lists the 3090 at 96.3 tok/s on Llama 3.1 8B Q4_K_M, against 59.0 tok/s for the RTX 5060 Ti. The 3090 matches the 5070 Ti on speed and clearly beats the 5060 Ti. Its one decisive edge over both is 8 GB of extra VRAM.
Hardware Corner's dedicated RTX 3090 page shows what that VRAM buys: 35.1 tok/s on Qwen3 32B at 4K context and 30.3 tok/s at 16K. In the dual-3060 thread, one commenter says a 32B model at Q4_K_M needs about 19 GB. That is more than any 16 GB card can hold without offloading.
The RTX 3090 vs RTX 4090 inference page covers the top of the range. Per Hardware Corner, the 4090 runs Qwen3 14B at 69.14 tok/s at 16K and the 5090 at 102.68 tok/s. The threads treat both as faster, not as better value.
The case for a new card
The new-card side gets a real hearing too:
- Price has moved against the 3090. In the 207 tok/s HN thread, one commenter says the cheapest working 3090 on eBay came to $1,800 CAD after exchange rate and shipping, against the roughly $700 they once sold for. In the dual-3060 thread, one reply calls 3090s overpriced and suggests an $800–1,000 budget for one. Another reply in the same thread says they have been $550–650 on r/hardwareswap. Asking prices clearly vary by venue and date, and the GPU Poet series is the dated reference here.
- Power and heat. NVIDIA rates the 3090 at 350 W board power with a 750 W system recommendation (spec page). It rates the RTX 5060 Ti at 180 W total graphics power (NVIDIA). In the 5080 + 3090 thread, a commenter says that kind of build draws 700 W at full load. Another caps each card at 220 W with nvidia-smi and reports losing only about 1 tok/s.
- Blackwell features. NVIDIA's RTX 5060 family page lists 5th-generation Tensor Cores and FP4 support on the 5060 Ti. The 3090 spec page lists 3rd-generation Tensor Cores. The counted threads mention FP8 and FP4 in passing but do not benchmark them, so this synthesis makes no throughput claim for FP4.
- 16 GB is enough for mixture-of-experts models. In an r/LocalLLaMA comparison of Qwen3.6 35B-A3B and Gemma 4 26B-A4B, the poster runs both on a 16 GB card at comparable speeds. Mixture-of-experts models are the main reason a 16 GB card no longer feels cramped.
- New 16 GB stock is not guaranteed. The HN thread on NVIDIA reportedly ending RTX 5070 Ti production (January 2026, with the 5060 Ti 16GB reportedly next) has commenters warning that 16 GB consumer cards may get scarce. That argues for buying new soon, if new is the plan.
Warranty comes up less than one might expect. None of the counted threads puts a figure on it. The practical point holds anyway: a new card comes with recourse, while a used 3090 bought from a stranger usually does not.
Used-buying risks the threads raise
- Mining history. The dual-3090 rig thread on HN argues both sides. One commenter warns that many used 3090s came off the crypto boom and may be overused. Another reports running seven ex-mining 3090s for two years without trouble. A third says two ex-mining cards bought about three years earlier have had no problems. In the 5080 + 3090 thread, a reply to someone who paid €390 for a second-hand 3090 asks whether it was a mining card.
- Corroded heatsinks. In the same dual-3090 thread, one commenter notes many eBay 3090 listings showing rusted or corroded heatsinks. That is a visible red flag worth checking in listing photos.
- Heat, noise and fit. The 2x3090 r/LocalLLaMA thread has a builder whose second 3090 would not fit in their case. The HN dual-3090 thread has commenters who needed a larger case or risers for airflow.
- Power estimates are contested. One HN commenter estimates 1,400 W for four 3090s at load. Another replies that four 3090s running inference on a large model draw closer to 350 W combined (thread). Actual draw depends heavily on workload and power limits.
None of the counted threads gives measured VRAM junction temperatures or before-and-after thermal-pad results. Treat repadding as a known Ampere maintenance item rather than something these sources quantify.
Dual RTX 3060 12GB vs a single 3090
The r/LocalLLM thread asking whether dual 12 GB 3060s are the right track is the clearest debate. The poster has a $400–600 budget. Most replies favour one 3090. One commenter says the 3060 is half as fast as a 3090 and that 2x12 GB yields less usable VRAM than one 24 GB card because of overhead. Another reports about 48 tok/s on Qwen2.5-14B Q5_K_M on a 3090, against about 19 tok/s on a friend's dual-3060 setup.
The dissent is about budget. One reply reports buying two 3060s for about $400 total from eBay and Facebook Marketplace. That commenter uses the pair for batch work and a 3090 Ti for interactive use, and says two 3060s will not generate tokens like a 3090-class card.
The 207 tok/s HN thread has the strongest pro-3060 voice. That commenter runs Qwen 27B and 35B on two 3060s at 8K–16K context, at 14 tok/s for dense models and 68 tok/s for mixture-of-experts. They say a three-card rig cost less than one 3090, and a reply challenges them to post a full parts list. In the r/LocalLLM budget thread, one reply puts a used RTX 3060 12GB at about $200.
Per Hardware Corner, a single RTX 3060 runs Qwen3 14B at 22.66 tok/s against 52.14 for the 3090. The parts list for the two-card route is in the dual RTX 3060 24GB build guide. What one 3060 can do alone is in the RTX 3060 12GB complete local LLM guide.
Who should buy what
- You need 27B–32B dense models at Q4 or better, on one card: used RTX 3090. It is the 24 GB card the threads keep coming back to. Inspect it for corrosion and mining wear, power-limit it, and budget against the current GPU Poet range rather than older forum prices. See the ASUS TUF RTX 3090 and ZOTAC RTX 3090 Trinity listings for new-stock pricing.
- You mostly run 8B–20B models or mixture-of-experts models, and want low power and a warranty: RTX 5060 Ti 16GB. It is slower than a 3090 in Hardware Corner's table, at 180 W per NVIDIA. See the ASUS Dual RTX 5060 Ti 16GB.
- You want 3090-class speed new, and 16 GB is enough: RTX 5070 Ti. Per Hardware Corner, it matches or beats the 3090 on 8B, 14B and gpt-oss 20B. See the MSI RTX 5070 Ti Ventus 3X.
- You have a hard budget around $400 and run batch jobs: dual RTX 3060 12GB. Accept roughly half the token rate. On whether a GPU is worth it at all for gpt-oss, see gpt-oss-20b on an RTX 3060 vs a Ryzen 5 5600G and gpt-oss-120b offload on an RTX 3060 vs a Ryzen 7 5800X. For a 3060 that also handles a media server, see Jellyfin NVENC plus a local LLM on one RTX 3060.
- Money is not the constraint: RTX 4090 or 5090. The threads treat both as faster, not better value.
Every thread counted
- we really all are going to make it, aren't we? 2x3090 setup. (r/LocalLLaMA, May 2026)
- On the right track with dual 12gb 3060's? (r/LocalLLM, July 2026)
- Is a 3050/60 a smart choice right now given the rumor that NVIDIA is rereleasing? (r/LocalLLM, March 2026)
- Best local ai for coding Nextjs project (r/LocalLLM, May 2026)
- Layman's comparison on Qwen3.6 35b-a3b and Gemma4 26b-a4b-it (r/LocalLLaMA, April 2026)
- RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8 (Hacker News, June 2026)
- We got 207 tok/s with Qwen3.5-27B on an RTX 3090 (Hacker News, April 2026)
- 25L Portable NV-linked Dual 3090 LLM Rig (Hacker News, September 2025; the submission links an r/LocalLLaMA build post)
- Nvidia Reportedly Ends GeForce RTX 5070 Ti Production, RTX 5060 Ti 16 GB Next (Hacker News, January 2026)
Tally across these nine: six are built on RTX 3090s or recommend one (1, 2, 4, 6, 7, 8). Three make the case for 16 GB or budget cards (3, 5, 9). Two describe 3090 used prices as inflated (2, 7). Two raise mining wear (6, 8).
Citations and sources
- https://gpupoet.com/gpu/learn/price/september-2026/nvidia-geforce-rtx-3090
- https://gpupoet.com/gpu/learn/price/february-2026/nvidia-geforce-rtx-3090
- https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090-3090ti/
- https://www.hardware-corner.net/gpu-ranking-local-llm/
- https://www.reddit.com/r/LocalLLaMA/comments/1tcf2dt/we_really_all_are_going_to_make_it_arent_we/
- https://www.reddit.com/r/LocalLLM/comments/1t9371m/best_local_ai_for_coding_nextjs_project/
- https://news.ycombinator.com/item?id=47838788
- https://news.ycombinator.com/item?id=48515454
- https://www.localscore.ai/accelerator/1
- https://www.localscore.ai/accelerator/860
- https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/
- https://www.reddit.com/r/LocalLLM/comments/1uv4rgo/on_the_right_track_with_dual_12gb_3060s/
- https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5060-family/
- https://www.reddit.com/r/LocalLLaMA/comments/1sqxiz0/laymans_comparison_on_qwen36_35ba3b_and_gemma4/
- https://news.ycombinator.com/item?id=46632955
- https://news.ycombinator.com/item?id=45300668
- https://www.reddit.com/r/LocalLLM/comments/1s8cymo/is_a_305060_a_smart_choice_right_now_given_the/
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.
