Skip to main content

Local LLM Hardware Digest — week of August 31, 2026

Issue of 2026-08-31 · sent to 38 subscribers

Local LLM Hardware Digest

Week of August 31, 2026

What moved this week in local-inference hardware: measured throughput, real price movement, and what the subreddits are arguing about. Every number links to the runs behind it.

Benchmarked this week
NVIDIA GeForce RTX 4070 Ti SUPER — 57.5 tok/s median at Q4
9 new runs from 1 source landed this week, on 16 GB of VRAM.
See every run and its source →
Budget GPU price movement
Quadro RTX 5000 down 19% to $627
Tracked at $775 before. Amazon pricing moves daily — the price at checkout is the one that counts.
View current price →
New this week
Mac Mini M4 vs RTX 3060 12GB for Local LLMs: Which Actually Wins?
Which of the two cheapest local-inference machines wins depends on model size, not brand. We price both builds and draw the line precisely.
Read it →
Best tok/s per dollar under $700
NVIDIA GeForce RTX 5060 Ti — 79.7 tok/s at 8B Q4 on 16 GB
Median of 8 community-reported runs. 16 GB is the tier where 8B models stop spilling to system RAM.
See the measurements →
Deciding rather than shopping?
The VRAM-tier table shows what each card fits at Q4 and how fast it actually generates.
Best GPUs for Running Local LLMs →

As an Amazon Associate, SpecPicks earns from qualifying purchases. Prices were accurate when this issue was generated and change frequently.