Skip to main content

Strix Halo LLM Benchmarks 2026: What AMD's Ryzen AI Max+ Delivers

Community llama.cpp benchmarks for AMD's Ryzen AI Max+ 395: gpt-oss-120b near 56 tok/s, dense 70B under 5 tok/s, and why bandwidth decides it.

Sept 2026 Strix Halo LLM benchmarks: gpt-oss-120b ~56 tok/s, Qwen3.6-35B ~63 tok/s, dense 70B under 5 tok/s. Why the 256GB/s bus decides it.

Strix Halo LLM Benchmarks 2026: What AMD's Ryzen AI Max+ Delivers

As of September 2026, a 128GB AMD Ryzen AI Max+ 395 ("Strix Halo") system runs OpenAI's gpt-oss-120b at about 55.6 tokens per second and Qwen3.6-35B-A3B at about 62.6 tok/s in llama.cpp's Vulkan backend, per the community-maintained Strix Halo Guide benchmark log. The same log records a dense Llama 3.1 70B at Q4_K_M at only 4.7–4.9 tok/s. The gap comes from the chip's roughly 256GB/s memory bus. Mixture-of-experts (MoE) models read only a few billion active parameters per token, so they run fast. Large dense models are limited by memory bandwidth and run slowly. Strix Halo's real advantage is capacity: 120B-class MoE models fit on one mini-PC, and no consumer discrete GPU can hold them.

What changed — September 2026 refresh

  • Measured decode figures added. The earlier version framed Strix Halo on memory capacity alone. The refresh adds a table of six community-measured llama.cpp decode figures on a 128GB Ryzen AI Max+ 395, including gpt-oss-120b at 55.57 tok/s and Qwen3.6-35B-A3B at 62.56 tok/s on Vulkan, per the Strix Halo Guide benchmark log. Each row also lists its Hugging Face weight size and a 256GB/s bandwidth ceiling.
  • The "70B fits" claim is now qualified with a speed. A dense Llama 3.1 70B at Q4_K_M fits in memory but decodes at only 4.7–4.9 tok/s, under its ~6 tok/s bandwidth ceiling, per the same Strix Halo Guide log.
  • The ROCm section was rewritten. The paragraph saying the ROCm gap "narrowed through 2025" is gone. In its place are measured results: Vulkan led decode at 97.73 vs 73.65 tok/s on Qwen3-Coder-30B-A3B, and ROCm led pp512 prefill at 1,344.65 vs 1,115.30 tok/s, per Soothill's August 2026 comparison. The section also notes kyuz0's stable ROCm 10.0 gfx1151 toolbox.
  • Links updated. The llama.cpp link now points to the current ggml-org/llama.cpp repository. A bare r/LocalLLaMA link that carried no figure was removed.

What "Strix Halo" Actually Is (and Isn't)

"Strix Halo" is AMD's codename for the Ryzen AI Max 300 series. It is a mobile/desktop-class APU that pairs Zen 5 CPU cores with an RDNA 3.5 integrated GPU (the Radeon 8060S on the top part, LLVM target gfx1151) and a unified LPDDR5X memory subsystem. AMD announced it at CES 2025. It ships in mini-PCs and high-end laptops such as the Framework Desktop, GMKtec EVO-X2, Beelink GTR9 Pro, and ASUS ROG Flow Z13. It is not in the same product tier as AMD's Instinct MI300X or Radeon Pro W7900. Those are datacenter and workstation-class discrete GPUs, sold separately from any host CPU and built for server-scale training and inference with their own HBM or GDDR memory. Coverage that mixes Strix Halo throughput figures with MI300X or W7900 numbers compares two different product categories, so this piece keeps them separate. For how the same chip handles gaming workloads, see SpecPicks' Strix Halo Gaming Benchmarks: What AMD's APU Can Really Do.

Why Strix Halo Matters for Local LLM Inference

For local LLM use, the key spec is memory capacity, not raw compute. Per AMD's published specifications, the top Ryzen AI Max+ 395 configuration has a 16-core/32-thread Zen 5 CPU, a 40-compute-unit RDNA 3.5 iGPU, and up to 128GB of on-package LPDDR5X-8000 memory on a 256-bit bus. The OS typically reserves about 2GB, which leaves well over 120GB as a shared pool for the GPU, NPU, and CPU. A discrete consumer GPU like the RTX 4090 tops out at 24GB of VRAM. That forces heavy quantization or model-splitting for anything much past a 30B-parameter model. A Strix Halo system can give the GPU a far larger share of memory. It can hold 70B dense weights and 100B-plus MoE weights entirely in memory, with no offloading to system RAM or disk. SpecPicks' Strix Halo vs. RTX 3060 12GB: Unified Memory or Discrete VRAM? and Best Mini PC for Local LLMs in 2026 cover that capacity-versus-bandwidth tradeoff in more depth.

What's Changed as of September 2026: Measured Strix Halo Numbers

The table below collects public community measurements of current open-weight models on a 128GB Ryzen AI Max+ 395. All rows are llama.cpp llama-bench token-generation (tg128) figures with no speculative decoding. File sizes are the summed weight files for the exact quant, taken from each Hugging Face repository.

The "bandwidth ceiling" column is a physics sanity check: 256GB/s divided by the bytes read per token. For MoE models, that is the file size scaled by active ÷ total parameters. Real decode speed always lands below this ceiling, and every row here does.

Model (quant)Weight file sizeActive / total paramsBandwidth ceiling (256GB/s)Measured decode, Vulkan RADVSource
Qwen3.6-35B-A3B (UD-Q4_K_M)22.1GB3B / 35B (MoE)~135 tok/s62.56 tok/sStrix Halo Guide
gpt-oss-120b (MXFP4)63.4GB5.1B / 117B (MoE)~92 tok/s55.57 tok/s (56.14 on llama.cpp b11146, re-checked Sept. 26, 2026)Strix Halo Guide
Qwen3.8-Flash-Next (UD-IQ4_XS)93.7GB6B / 125B (MoE, plus n-gram embeddings and MTP head)~57 tok/s (conservative)27.16 tok/s (28.61 on b11146)Strix Halo Guide
Qwen3.5-122B-A10B (UD-Q4_K_XL)77.0GB10B / 122B (MoE)~41 tok/s22.9 tok/s (ROCm: 21.3)Soothill, Aug. 2026
Nemotron 3 Super 120B-A12B (UD-IQ4_XS)64.5GB12B / 120B (MoE)~40 tok/s18.43 tok/sStrix Halo Guide
Llama 3.1 70B Instruct (Q4_K_M, dense; Ollama Vulkan)42.5GB70B / 70B (dense)~6 tok/s4.7–4.9 tok/sStrix Halo Guide

How to read it:

  • Active parameters, not total size, set decode speed. The 22GB Qwen3.6-35B-A3B runs about 13× faster than the 42GB dense Llama 3.1 70B because it reads roughly 2GB of weights per token instead of 42GB.
  • 120B-class models are practical, not fast. gpt-oss-120b is the standout because only 5.1B parameters are active per token. Per the Strix Halo Guide, its prompt processing also held at 293.73 tok/s even at a 65,536-token prompt. Nemotron 3 Super and Qwen3.5-122B, with 10–12B active parameters, land around 18–23 tok/s: usable for chat, slow for agent loops.
  • The largest models barely fit. The 93.7GB Qwen3.8-Flash-Next file loads within the 128GB pool, but the guide notes that its September 26 re-check ran last and filled swap. Treat models near 90GB+ as tight fits that leave little room for long-context KV cache.
  • Numbers move with the software stack. The guide's September 26 re-check on llama.cpp b11146 raised prompt processing 13–28% over the same-night original builds, while decode moved only a few percent. Two boxes on different kernels, Mesa versions, or llama.cpp builds can disagree by a few tok/s. The guide's README treats about 2% tg128 spread as normal between well-matched systems.

Throughput also falls as context fills: per the same guide, Qwen3.6-35B-A3B decodes at 32.23 tok/s after a filled 128K KV cache, versus about 62 tok/s at short context.

The Real Bottleneck: Memory Bandwidth, Not Compute

Community benchmark logs such as the Strix Halo Guide and discussion on the llama.cpp GitHub repository consistently identify memory bandwidth, not raw TFLOPs, as the binding limit on Strix Halo's token-generation ("decode") speed. Autoregressive decoding is largely memory-bandwidth-bound, because each new token requires re-reading the active weights from memory. AMD's published ~256GB/s figure for the 256-bit LPDDR5X-8000 bus is far below the roughly 1TB/s of an RTX 4090's GDDR6X and the multi-terabyte-per-second HBM3 of datacenter parts like the MI300X. The dense 70B row above shows the result: 42.5GB of weights over a 256GB/s bus caps decode near 6 tok/s before any software overhead, and the measured 4.7–4.9 tok/s sits just under that cap. Strix Halo gives up raw decode speed in exchange for running models that won't fit on a VRAM-limited discrete card at all. That is a tradeoff, not an unqualified win. Prompt processing (prefill) is more compute-bound and scales better: the guide's gpt-oss-120b run measured 726.99 tok/s pp512 against 55.57 tok/s decode.

Software Maturity: ROCm vs. Vulkan

The early Strix Halo LLM community leaned on llama.cpp's Vulkan backend because AMD's ROCm stack had historically supported consumer RDNA hardware less well than Instinct-class GPUs. ROCm has since shipped native gfx1151 support. kyuz0's widely used Strix Halo llama.cpp toolboxes now list a stable ROCm 10.0 container alongside the Vulkan RADV one, with Vulkan RADV still described as the most stable and compatible choice for most users.

The measured picture as of 2026 is a split by workload, not a single winner:

  • Vulkan usually wins decode. In Soothill's August 2026 comparison, Qwen3-Coder-30B-A3B Q4_K_S generated 97.73 tok/s on Vulkan versus 73.65 tok/s on ROCm.
  • ROCm usually wins prompt processing. The same comparison measured ROCm ahead on pp512, at 1,344.65 versus 1,115.30 tok/s for Qwen3-Coder and 339.9 versus 285.0 tok/s for Qwen3.5-122B-A10B. The Strix Halo Guide found the same pattern: HIP was 24.8% faster at pp16384 on Qwen3.6-35B-A3B, while Vulkan was 18.1% faster at tg128.

In practice, use Vulkan for interactive chat and coding. Test ROCm for RAG ingestion, long-document summarization, and batch serving, where prefill dominates. Either way, verify current driver and backend support for your specific model and quantization before buying rather than assuming parity with Nvidia's CUDA ecosystem. kyuz0's toolbox notes, for example, still recommend flash attention and --no-mmap on Strix Halo to avoid crashes and slowdowns. SpecPicks' Ryzen AI Halo mini-PC coverage and DGX Spark rival review roundup track this in more detail.

Strix Halo vs. Discrete GPU Options

PlatformMemory type / capacityPeak memory bandwidth (mfr. spec)Typical form factor
Ryzen AI Max+ 395 ("Strix Halo")Up to 128GB unified LPDDR5X-8000~256GB/sMini-PC / high-end laptop
RTX 409024GB GDDR6X~1,008GB/sDiscrete desktop GPU
RTX 509032GB GDDR7~1,792GB/sDiscrete desktop GPU
Instinct MI300X192GB HBM3Multi-TB/s classDatacenter accelerator (not consumer)

Figures come from each vendor's published product specifications, not independent testing. Sustained throughput depends heavily on software stack, quantization, and workload. A model that fits in 24–32GB of VRAM, such as the 22GB Qwen3.6-35B-A3B quant, has a far higher bandwidth ceiling on an RTX 4090 or 5090; the 63GB gpt-oss-120b doesn't fit on either card without offloading. For discrete-GPU context in the same class of AI workload, see RTX 5090 Benchmark 2026: 4K Gaming & AI vs. AMD and RTX 5090 AI Desktop: 2026 Build Guide & Benchmarks.

Building an AMD-Based AI Rig vs. Buying a Strix Halo Mini-PC

The unified-memory architecture is currently exclusive to the mobile/mini-PC Ryzen AI Max lineup. Readers considering a DIY AMD desktop instead should know that a standard AM4/AM5 build still pairs a conventional motherboard, such as the ASUS ROG Strix B550-F Gaming Motherboard ($252.19), with a separate discrete GPU rather than an integrated unified-memory chip. Fast local storage matters on either platform. Weight files for the models above run from 22GB to over 90GB and must be read from disk at startup and during model swaps. An enclosure like the ASUS ROG Strix Arion USB 3.2 M.2 NVMe Enclosure ($62.28) speeds that up for external or swappable storage. Long inference sessions also generate meaningful heat, so case airflow is a reasonable line item for a dedicated AI-rig build. The Cooler Master MasterFan MF120 Halo 3-in-1 ARGB fans ($108.96) are one option. Prices reflect the listing at time of publication and may vary, so check the current price before purchasing.

Cost-Performance Takeaways

  • Strix Halo's advantage is capacity per dollar for large quantized models that don't fit in 16–32GB of discrete VRAM, not raw tokens per second.
  • Pick MoE models: per the community figures above, active-parameter count, not total size, decides decode speed.
  • Discrete GPUs with less memory but far more bandwidth (RTX 4090, RTX 5090) still generally win on decode speed for models that fit in their VRAM.
  • Datacenter parts like the MI300X aren't a relevant price/performance comparison for a home or small-business local-LLM rig. They are server accelerators, not desktop or mini-PC components.
  • Backend choice matters: Vulkan for generation-heavy use and ROCm for prompt-heavy use, per the public comparisons above. Verify current support for your target model and quantization before buying.

Bottom Line

"Strix Halo benchmark LLM" searches often turn up a mix of consumer-APU and datacenter-GPU numbers that don't belong in the same comparison. Judged on its own terms, the Ryzen AI Max+ 395's pitch for local LLM work is unified memory capacity rather than raw throughput. As of September 2026, public community benchmarks show it running 120B-class MoE models such as gpt-oss-120b at about 55 tok/s, a class of model no single consumer discrete card can hold. Dense 70B models remain bandwidth-bound at under 5 tok/s. ROCm now has a stable gfx1151 release, but Vulkan remains the faster decode path in public comparisons. For broader context on how fast the open-weight model landscape is moving around hardware like this, see Open-Weight Models Caught Up on Cyber Benchmarks in Four Months.

Citations and sources

  • https://github.com/hogeheer499-commits/strix-halo-guide/blob/main/BENCHMARKS.md
  • https://github.com/hogeheer499-commits/strix-halo-guide
  • https://www.soothill.io/blog/2026/08/03/llamacpp-vulkan-vs-rocm-strix-halo/
  • https://github.com/kyuz0/amd-strix-halo-toolboxes
  • https://github.com/ggml-org/llama.cpp
  • https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
  • https://huggingface.co/ggml-org/gpt-oss-120b-GGUF
  • https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
  • https://huggingface.co/unsloth/Qwen3.5-122B-A10B-GGUF
  • https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF
  • https://huggingface.co/bartowski/Meta-Llama-3.1-70B-Instruct-GGUF
  • https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html
  • https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
  • https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
  • https://frame.work/products/desktop-diy-amd-aimax300

This piece is editorial synthesis based on publicly available information. No independent first-party benchmarking is reported.

Products mentioned in this article

Amazon & eBay listings, plus full specs and alternatives on each product page.

As an Amazon Associate, SpecPicks earns from qualifying purchases; we also earn on qualifying eBay purchases via the eBay Partner Network. Prices shown were last tracked at crawl time and may vary — check the listing for the current price.

Frequently asked questions

How fast does Strix Halo run LLMs as of September 2026?
Per the community-maintained Strix Halo Guide benchmark log, a 128GB Ryzen AI Max+ 395 running llama.cpp's Vulkan backend generates about 62.6 tokens per second on Qwen3.6-35B-A3B (UD-Q4_K_M) and about 55.6 tokens per second on gpt-oss-120b (MXFP4). Larger-active MoE models such as Nemotron 3 Super 120B-A12B land around 18 tokens per second, and a dense Llama 3.1 70B at Q4_K_M manages only 4.7 to 4.9 tokens per second.
Why are dense 70B models so slow on Strix Halo?
Token generation is limited by memory bandwidth. Strix Halo's 256-bit LPDDR5X-8000 bus peaks at roughly 256GB/s, and a 42.5GB Q4_K_M 70B model must be read in full for every token, which caps decode near 6 tokens per second before software overhead. Mixture-of-experts models read only their active parameters per token, so they run many times faster.
Is Strix Halo faster than an RTX 4090 for running LLMs locally?
Not on raw decode speed for models that fit in 24GB of VRAM. The RTX 4090's GDDR6X delivers about 1,008GB/s, roughly four times Strix Halo's bus. Strix Halo's advantage is capacity: a 63GB model such as gpt-oss-120b fits entirely in its unified memory but does not fit on a 24GB card without offloading.
Should I use ROCm or Vulkan for LLM inference on Strix Halo?
Public comparisons show a split by workload. In an August 2026 comparison published by Soothill, Vulkan generated 97.73 tokens per second on Qwen3-Coder-30B-A3B versus 73.65 on ROCm, while ROCm led prompt processing at 1,344.65 versus 1,115.30 tokens per second. Vulkan is the usual choice for chat and coding; ROCm is worth testing for long-prompt, RAG and batch workloads.
Is Strix Halo the same as AMD's MI300X for LLM workloads?
No. Strix Halo (Ryzen AI Max 300 series) is a mobile and mini-PC APU with up to 128GB of unified LPDDR5X memory, while the MI300X is a datacenter accelerator with 192GB of its own HBM3 memory, sold as a server part. Benchmark figures for one do not transfer to the other.

— Mike Perry · Updated 2026-10-11

Parts this article names

Amazon Associate — prices tracked 2026-10-10, may vary.