Local LLM Hardware Digest — week of September 7, 2026
Local LLM Hardware Digest
Week of September 7, 2026
What moved this week in local-inference hardware: measured throughput, real price movement, and what the subreddits are arguing about. Every number links to the runs behind it.
|
Benchmarked this week
NVIDIA GeForce RTX 3060 — 42 tok/s median at Q4
12 new runs from 4 sources landed this week, on 12 GB of VRAM.
See every run and its source →
|
|
Budget GPU price movement
Quadro RTX 4000 down 11% to $280
Tracked at $314 before. Amazon pricing moves daily — the price at checkout is the one that counts.
View current price →
|
|
New this week
RTX 5070 Ti vs RTX 5090 for Local LLMs: 16GB vs 32GB
The RTX 5090's 32 GB is a memory ceiling, not a speed grade. Which models fit 16 GB, which need 32 GB, and where the KV cache breaks both.
Read it →
|
|
Best tok/s per dollar under $700
NVIDIA GeForce RTX 5060 Ti — 79.7 tok/s at 8B Q4 on 16 GB
Median of 8 community-reported runs. 16 GB is the tier where 8B models stop spilling to system RAM.
See the measurements →
|
|
On the radar
Intel Core I5 13400F — 39 mentions and counting
Turning up across r/LocalLLaMA and r/hardware. Not yet in the SpecPicks benchmark catalog — it is in the queue, and this is a mention count, not a verdict.
Read the thread →
|
|
Deciding rather than shopping?
The VRAM-tier table shows what each card fits at Q4 and how fast it actually generates.
Best GPUs for Running Local LLMs →
|
As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices were accurate when this issue was generated and change frequently.