Why the sub-4B tier needs its own buying guide
Most local-LLM buying advice starts at 8B parameters and asks how much VRAM you can afford. This guide covers the tier below that: Qwen3 0.6B through 4B, Llama 3.2 1B and 3B, Gemma 3 1B, and Phi-4-mini. These models are small enough that capacity stops being the question. A 3B model at Q4_K_M is a 1.87 GiB file per Geerling's llama.cpp output. Qwen3 0.6B at Q4_K_M is 0.40 GB per Unsloth's GGUF repository. Everything from a Raspberry Pi to a gaming PC can hold them.
What separates hardware at this tier is two other numbers. The first is memory bandwidth, because token generation streams the whole weight file once per token. The second is watts, because sub-4B models usually end up doing always-on background work (voice-intent routing, log triage, document tagging, home automation) where the box idles far more than it computes.
That changes the shape of a good purchase. A discrete GPU is by far the fastest option, but it's also the hungriest and the most expensive to leave running. A Raspberry Pi is slow, but cheap to run and tiny. A desktop APU sits between them. The right pick depends on whether a person is waiting on each answer and whether you'll want bigger models later.
This guide ranks five parts from the SpecPicks catalog against public measurements from Geerling's ai-benchmarks project, a published single-board-computer LLM study, Zen 3 llama.cpp results from TechHara, and vendor spec pages. Where no source measured a specific part, the figure is scaled from the nearest measurement and labelled as an estimate. The winner, for most readers, is the card that looks like overkill, for one reason: it's the only pick here that doesn't need replacing when a 3B model stops being enough.
Step 0: diagnose before you buy
Answer three questions before choosing anything. They send you to different picks.
1. Is the box always on, or used in bursts? If it runs 24/7 and mostly waits (a Home Assistant intent parser, a nightly summarizer), idle and peak watts dominate the cost of ownership. Go to the Raspberry Pi 4 8GB or the Vilros kit. If you use it in bursts from your desk, power barely matters. Go to the RTX 3060 12GB.
2. Is a person waiting on each response? Interactive chat and voice assistants with a real system prompt need fast prefill as well as generation. On a Raspberry Pi 4, a 300-token prompt can take tens of seconds before the first token appears (estimates below). If someone waits, get the Ryzen 5 5600G at minimum, and the RTX 3060 12GB ideally. If nobody waits, the Pi is fine.
3. Will you move to 8B-14B within a year? If yes, buy for that tier now. The RTX 3060 12GB runs Qwen3 14B at 31.2 tok/s per Hardware Corner, while every other pick here falls to single digits or can't load it. If the plan is "a small model does one job forever", the cheaper picks are the right call.
If you run other services on the same machine (a NAS, a media server, containers), the Ryzen 7 5800X is the host that keeps prompt processing responsive while those services compete for cores.
Comparison table
| Pick | Best for | Key spec | Price range (catalog, 2026-09-16, may vary) | Verdict |
|---|---|---|---|---|
| 🏆 MSI Gaming GeForce RTX 3060 12GB | Interactive use, upgrade path to 14B | 12 GB GDDR6, 360 GB/s | ~$230 average market (Hardware Corner) to $479.99 catalog | 6-77× faster than the other picks on 3B; needs a host PC |
| 💰 AMD Ryzen 5 5600G | GPU-free desktop server | 6 Zen 3 cores + Radeon iGPU, 65 W | $199.99 | ~19 tok/s on 3B (est.); cheapest interactive-capable box |
| 🎯 Raspberry Pi 4 Model B 8GB | Always-on, low-watt edge host | 4× Cortex-A72, 8 W inference peak | $178.81 | Fine for 0.5B-1B routing; too slow for 3B chat |
| ⚡ AMD Ryzen 7 5800X | Shared host running other services | 8 Zen 3 cores / 16 threads, 105 W | $254.90 | More prefill headroom; needs a discrete GPU to boot |
| 🧪 Vilros Raspberry Pi 4 8GB Basic Starter Kit | First build, complete in one box | Pi 4 8GB + fan-cooled case + USB-C PSU | $211.99 | Cheapest complete, cooled Pi setup |
🏆 Best Overall: MSI Gaming GeForce RTX 3060 12GB
The MSI Gaming GeForce RTX 3060 12GB
Spec chips: 12 GB GDDR6 · 192-bit bus · 360 GB/s bandwidth · 3,584 CUDA cores · 170 W card · 550 W PSU required (NVIDIA, Hardware Corner)
Pros
- 122.85 tok/s generation on Llama 3.2 3B Q4_K_M and 323.07 tok/s on the 1.1B TinyLlama (Geerling #40).
- Prefill of 2,800.87 tok/s at 4,096 tokens on the 3B model, so long prompts start answering almost instantly.
- 12 GB carries you to Qwen3 14B at 31.2 tok/s (4K context) with no second purchase (Hardware Corner).
- The host barely matters: a Raspberry Pi CM5 driving the card reached 112.77 tok/s on the same 3B model (Geerling #40).
Cons
- Needs a host PC with a PCIe slot and a 550 W supply.
- It's the highest-power pick. Geerling measured a 214 W system peak on the 3B run.
- Catalog listings for this card vary widely; the MSI listing was $479.99 on 2026-09-16, against the ~$230 average market value Hardware Corner reports.
For sub-4B models, the RTX 3060 12GB is overkill at first glance, and that's the argument for it. At this tier, generation is limited by how fast weights move. The card's 360 GB/s of GDDR6 bandwidth is roughly seven times a dual-channel DDR4-3200 desktop's ~51.2 GB/s. In Geerling's measurements, Llama 3.2 3B generated at 122.85 tok/s on a Core Ultra 265K host. That's more than five times the 23.81 tok/s of a Ryzen AI 5 340 laptop CPU and about 77 times the 1.60 tok/s of a Pi 400 (Geerling README).
Prefill matters even more for interactive use. At 2,800.87 tok/s on a 4,096-token prompt, a Home Assistant request carrying a long device list starts answering in about a second. On CPU-only hardware, the same request takes seconds to minutes.
The real case for the card is what happens next. Small models route and classify well, but they hit a quality ceiling on reasoning and retrieval. When that happens, the RTX 3060 moves up to Qwen3 8B at 55.2 tok/s and Qwen3 14B at 31.2 tok/s (Hardware Corner) without a second purchase. The alternative is buying a Pi or an APU box now and a GPU later.
Price-freshness note: prices come from the retailer at the last refresh and may vary. Confirm at checkout. See Full Details →
💰 Best Value: AMD Ryzen 5 5600G
Spec chips: 6 Zen 3 cores / 12 threads · 7-core Radeon iGPU at 1,900 MHz · dual-channel DDR4 up to 3200 MT/s · PCIe 3.0 · 65 W default TDP (AMD)
Pros
- Boots without a graphics card, so one chip builds a complete low-cost server.
- Estimated ~19 tok/s on 3B Q4_K_M and ~15 tok/s on Qwen3 4B, which is interactive-grade for short answers.
- The integrated Radeon doubles prompt processing over the CPU cores (TechHara).
- Its PCIe slot can take a GPU later.
Cons
- Dual-channel DDR4 caps generation, and faster cores don't raise that ceiling.
- A single memory stick halves the bandwidth, so budget for a matched pair.
- Still needs a motherboard, RAM, PSU and case, so the full system costs more than the chip.
The 5600G is the cheapest way into genuinely usable sub-4B inference without a graphics card. No source has benchmarked it on a 3B model directly, but two Zen 3 measurements bracket it well. TechHara's six-core Ryzen 5 5600H generated ~10 tok/s on a 3.56 GiB 7B file, with ~34 tok/s of CPU prefill (TechHara). A Ryzen 7 5700G reported 8.8 tok/s on Mistral 7B, CPU-only (ROCm #2774). Scaling by weight size to a 1.87 GiB 3B file gives about 19 tok/s, under the bandwidth ceiling of 51.2 GB/s divided by the file size.
The iGPU is worth switching on for prefill. TechHara saw prompt processing roughly double, from ~34 to ~76 tok/s, on Vega graphics via Vulkan, while generation "remained nearly identical (~10 t/s)". The iGPU shares the same memory bus, so it hits the same generation ceiling. It still cuts the wait before the first token in half.
llama.cpp developer Johannes Gäßler found that "just 5 threads are enough to fully utilize the memory bandwidth provided by dual channel memory" (Gäßler). On a 5600G, set -t 5 or -t 6 and leave the other threads for other services.
Price-freshness note: prices come from the retailer at the last refresh and may vary. Confirm at checkout. See Full Details →
🎯 Best for Always-On Edge Hosts: Raspberry Pi 4 Model B 8GB
The Raspberry Pi 4 Model B 8GB
Spec chips: Broadcom BCM2711, 4× Cortex-A72 @ 1.8 GHz · 8 GB LPDDR4 · 2× USB 3.0 · 5 V / 3 A USB-C · 8 W peak during LLM inference (Pi 4 product brief, arXiv 2511.07425)
Pros
- Measured 6.44 tok/s on Qwen2 0.5B and 24.09 tok/s on SmolLM 135M (lemonade-benchmark Pi 4 results).
- An 8 W inference peak makes it cheap to leave running.
- Silent with a small fan, and it fits anywhere.
Cons
- Slow on 3B models: 1.60 tok/s on the same BCM2711 chip in a Pi 400 (Geerling README).
- Prefill is the real weakness, so long system prompts mean long waits.
- 8 GB soldered, no PCIe slot, and no upgrade path.
The Pi 4 earns its place for one job: a low-watt box that runs a sub-1B model around the clock for work nobody watches. The study behind arXiv 2511.07425 concluded that the Pi 4 suits "ultra-lightweight tasks (e.g., command parsing)". Its raw results show why: 6.44 tok/s on Qwen2 0.5B, 4.76 on TinyLlama 1.1B, and 2.04 on Llama 3.2 1B. Scaled by file size, Qwen3 0.6B at Q4_K_M should run near 5.7 tok/s, which is enough to emit a 20-token intent JSON in under four seconds.
Two practical requirements come from the Pi 4 datasheet. The firmware throttles the CPU so the chip "never exceeds 85 degrees C", so sustained inference needs active cooling. And the SD slot peaks at 50 MB/s, which means a 24/7 host should boot from a USB 3.0 SSD. For a model this small, the reason is card wear from constant log writes, not load speed.
If you're buying new specifically for LLM work, note that the Pi 5 is substantially faster: 4.61 tok/s on Llama 3.2 3B against the Pi 400's 1.60 (Geerling README). The Pi 4 pick is for readers who have one, or whose job fits under 1B.
Price-freshness note: prices come from the retailer at the last refresh and may vary. Confirm at checkout. See Full Details →
⚡ Best Performance: AMD Ryzen 7 5800X
Spec chips: 8 Zen 3 cores / 16 threads · 105 W default TDP · DDR4 up to 3200 MT/s · PCIe 4.0 · discrete graphics card required (AMD)
Pros
- Two more cores than the 5600G, so CPU prefill runs up to about a third faster (estimate).
- Enough threads to run a small model alongside a NAS, media server or containers without starving either.
- PCIe 4.0, and a natural host for an RTX 3060 12GB.
Cons
- Generation is no faster than the 5600G. It uses the same dual-channel DDR4 bus.
- No integrated graphics. AMD lists "Discrete Graphics Card Required".
- 105 W TDP, and the listing notes "Cooler not included".
The 5800X is the "performance" pick in a specific sense: it's the CPU host that keeps prompt processing responsive when the small model isn't the only thing on the machine. Prefill is compute-bound, so its eight cores beat the 5600G's six. Generation is bandwidth-bound, so it doesn't. Gäßler's measurements on a Zen 2 system with dual-channel DDR4-3200 found that more threads than physical cores "can actually be detrimental" (Gäßler). The 5800X's extra threads are there for your other services, not for tokens per second.
Its best use is as the host for the Best Overall pick. With the RTX 3060 holding the model, the CPU mostly idles during generation, and the 5800X's cores stay free for everything else the box does. Since it can't boot without a graphics card, a 5800X build that never gets a GPU is the wrong purchase. In that case, buy the 5600G.
Price-freshness note: prices come from the retailer at the last refresh and may vary. Confirm at checkout. See Full Details →
🧪 Budget Pick: Vilros Raspberry Pi 4 8GB Basic Starter Kit
The Vilros Raspberry Pi 4 8GB Basic Starter Kit
Spec chips: Pi 4 Model B 8GB · aluminum alloy case with pre-installed fan · USB-C power supply with on/off switch (per listing)
Pros
- Covers the two datasheet requirements, 5 V / 3 A power and active cooling, in one purchase.
- At $211.99 against $178.81 for the bare board (catalog, 2026-09-16), the case, fan and PSU add about $33.
- Same measured performance as the bare Pi 4.
Cons
- Still needs boot storage. Add a USB 3.0 SSD for a 24/7 host.
- Same 3B-and-up performance limits as any Pi 4.
- A Pi 5 kit is the better starting point if you're buying new specifically for LLM speed.
For a first build, the kit is the lowest-risk way to get a Pi running inference reliably. The two most common failures on always-on Pi hosts are underpowered supplies, which cause brownouts that look like software crashes, and thermal throttling during sustained generation. The Pi 4 datasheet calls for a supply "capable of delivering 5V at 3A". The kit's listing describes a "Fan Cooled Heavy Duty Aluminum Alloy Case" and "a USB-C Raspberry Pi 4 compatible power supply". That handles both failures before they happen.
Performance matches the bare board: about 5.7 tok/s estimated on Qwen3 0.6B Q4_K_M, at an 8 W measured peak for the Pi 4 (arXiv 2511.07425). It's a budget pick in total-cost terms. Buying the case, fan and a proper supply separately usually costs more than the difference.
Price-freshness note: prices come from the retailer at the last refresh and may vary. Confirm at checkout. See Full Details →
What to look for in small-language-model hardware
Memory bandwidth over core count
Generation reads the full weight file once per token, so tokens per second is capped by bandwidth divided by file size. Dual-channel DDR4-3200 delivers about 51.2 GB/s, and the RTX 3060's GDDR6 delivers 360 GB/s (Hardware Corner). Extra CPU cores raise prefill, not generation.
Idle power and 24/7 cost
Sub-4B hosts mostly wait. At $0.15/kWh, every 10 W of continuous draw costs about $13 a year. Measure the whole system at the wall before trusting a spec-sheet TDP.
Storage for the model library
Once loaded, the model lives in RAM, so storage speed only affects cold starts. Reliability is what matters: always-on boxes write logs constantly, and microSD cards wear out under that load.
Thermal headroom and sustained clocks
Inference holds every core at full load for as long as a response takes. The Pi 4 throttles to stay under 85°C (datasheet), and the 5800X ships without a cooler. Budget for cooling up front.
Upgrade path to the 8B-14B tier
Only a 12 GB GPU runs 14B models interactively. The Pi can't take one at all, and the AM4 picks can add one later. Decide now whether the box is a dead end or a first step.
Benchmark table: sub-4B models across platforms
| Model (Q4 unless noted) | Platform | Generation | Prefill | Peak power | Source |
|---|---|---|---|---|---|
| TinyLlama 1.1B Q4_K_M | RTX 3060 12GB (Core Ultra 265K host) | 323.07 tok/s | 4,394.09 tok/s (pp4096) | 202 W system | Geerling #40 |
| Llama 3.2 3B Q4_K_M | RTX 3060 12GB (Core Ultra 265K host) | 122.85 tok/s | 2,800.87 tok/s (pp4096) | 214 W system | Geerling #40 |
| Llama 3.2 3B Q4_K_M | RTX 3060 12GB (Pi CM5 host) | 112.77 tok/s | 2,531.08 tok/s (pp4096) | 192.3 W system | Geerling #40 |
| Llama 3.2 3B | Ryzen AI 5 340 laptop CPU | 23.81 tok/s | not reported | 51.1 W | Geerling README |
| Llama 3.2 3B | Ryzen 5 5600G CPU | ~19 tok/s (estimate) | ~70 tok/s (estimate) | 65 W TDP | Scaled from TechHara |
| Llama 3.2 3B | Intel N150 mini PC | 9.06 tok/s | not reported | 26.4 W | Geerling README |
| Llama 3.2 3B | Raspberry Pi 5 8GB | 4.61 tok/s | not reported | 13.9 W | Geerling README |
| Llama 3.2 3B | Raspberry Pi 400 (Pi 4's BCM2711) | 1.60 tok/s | not reported | 6 W | Geerling README |
| Qwen2 0.5B Q4_0 | Raspberry Pi 4 | 6.44 tok/s | not reported | 8 W (study peak) | lemonade-benchmark, arXiv |
| Llama 3.2 1B | Raspberry Pi 4 | 2.04 tok/s | not reported | 8 W (study peak) | same |
The 5600G prefill estimate scales TechHara's ~34 tok/s CPU pp512 on a 6.74B model to a 3.21B model. Enabling the iGPU should roughly double it. Two more figures follow from the table. First, the host CPU barely changes GPU results: a Pi CM5 host lands within about 8% of a Core Ultra 265K on the 3B run. Second, energy per token favours the GPU at full load: 214 W at 122.85 tok/s is about 0.48 kWh per million tokens, against about 1.0 kWh for the Pi 400's 6 W at 1.60 tok/s.
Frequently asked questions
Do I need a discrete GPU for models under 4B parameters?
No. A 3B model at Q4_K_M is about 1.87 GiB, and Qwen3 0.6B is 0.40 GB, so a Raspberry Pi, an APU desktop or any gaming PC can hold them in memory. A discrete card buys speed, not capability: the RTX 3060 12GB measured 122.85 tok/s on Llama 3.2 3B against 1.60 tok/s on a Pi 400. Most builders who choose one do it for the upgrade path to 8B-14B models.
How much RAM should a small-model host have?
For the APU and CPU picks, 32 GB in a matched dual-channel kit is a comfortable target and 16 GB is enough for sub-4B models alone. The model is small, but the cache, the operating system and any other services share the same memory. More important than capacity is running two matched sticks: a single-stick setup halves memory bandwidth, and with it the generation ceiling.
Does storage speed affect inference performance?
Only at load time. Once the model sits in RAM, the drive is idle during generation. Loading is quick even on slow media: the Pi 4's 50 MB/s SD slot reads a 0.40 GB model in about eight seconds. What storage really affects on always-on boxes is reliability. Constant log writes wear out microSD cards, so boot a 24/7 Pi or mini server from an SSD.
What is the real annual power cost of a 24/7 local model host?
It depends mostly on idle draw, because a small-model server spends most of its life waiting. At $0.15 per kWh, every 10 W of continuous draw costs about $13 a year. A Raspberry Pi 4 peaked at 8 W during inference in a published SBC study, while a desktop with a discrete GPU can exceed 200 W under load. Measure your own system at the wall before comparing.
When should I skip this tier and buy 12GB of VRAM instead?
If the plan includes coding assistance, retrieval over real documents or multi-turn conversation, buy the 12 GB card now. Sub-4B models route and classify well but lose ground on reasoning-heavy work. The RTX 3060 12GB runs Qwen3 8B at 55.2 tok/s and Qwen3 14B at 31.2 tok/s at 4K context per Hardware Corner, so it covers the next tier without another purchase.
Related guides
- Best budget GPU for local LLMs
- Best mini PC for local LLMs
- Raspberry Pi 4 8GB for local LLMs
- LLM VRAM requirements by model
— Mike Perry · Last verified 2026-09-16
Citations and sources
- Jeff Geerling, ai-benchmarks issue #40: RTX 3060 results (accessed 2026-09-16)
- Jeff Geerling, ai-benchmarks README (accessed 2026-09-16)
- Hardware Corner, RTX 3060 12GB LLM benchmarks (accessed 2026-09-16)
- Nguyen & Nguyen, "An Evaluation of LLMs Inference on Popular Single-board Computers", arXiv 2511.07425 (accessed 2026-09-16)
- lemonade-benchmark, Raspberry Pi 4 Ollama results CSV (accessed 2026-09-16)
- TechHara, llama.cpp benchmark: CPU vs iGPU on Ryzen 5 5600H (accessed 2026-09-16)
- ROCm issue #2774: APU support for 5600G/5700G (accessed 2026-09-16)
- Johannes Gäßler, llama.cpp performance testing (accessed 2026-09-16)
- Unsloth, Qwen3-0.6B-GGUF (accessed 2026-09-16)
- NVIDIA, GeForce RTX 3060 family specifications (accessed 2026-09-16)
- AMD, Ryzen 5 5600G specifications (accessed 2026-09-16)
- AMD, Ryzen 7 5800X specifications (accessed 2026-09-16)
- Raspberry Pi 4 Model B datasheet (accessed 2026-09-16)
- Raspberry Pi 4 Model B product brief (accessed 2026-09-16)
This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.
