NVIDIA B200 — Benchmarks & Specs
Bottom line: how fast is the NVIDIA B200?
In a batched serving run it reaches 12841.0 tokens/sec aggregate throughput (many concurrent requests, not one user's generation speed) running llama2:70b at FP4 under trt-llm, per Lambda AI. In MLPerf Inference v5.1 — Llama 2 70B Offline it scores 102,725 tok/s, per Lambda AI. It has 192 GB of VRAM, and the largest model on file at 4-bit quantization on this card is gpt-oss-120b (117B parameters, MXFP4), generating 60000.0 tokens/sec under trt-llm, per NVIDIA Developer Blog. Larger models have no sourced run on file.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA B200 is a datacenter accelerator from the Blackwell family released in 2024 from NVIDIA. Key on-paper specs include 192 GB of HBM3e memory, 8000 GB/s of memory bandwidth, 1000W TDP. Data on this page draws on 3 synthetic benchmark results, 6 community AI inference reports, compiled from NVIDIA Developer Blog, Nebius / MLCommons MLPerf Inference v6.0, Data Science Collective / Cloudrift AI, Lambda AI, NVIDIA Blog; each table row links to its source. Read this page when shopping the NVIDIA B200, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| gpt-oss-120b | MXFP4 trt-llm batched | 60000.0 tok/s aggregate | — | NVIDIA Developer Blog (SemiAnalysis InferenceMAX v1) 2025-10-09 | |
| gpt-oss-120b | NVFP4 tgi batched | 30000.0 tok/s aggregate | — | NVIDIA Developer Blog / SemiAnalysis InferenceMAX 2025-10-13 | |
| llama2:70b | FP4 trt-llm batched | 11264.0 tok/s aggregate | — | NVIDIA Developer Blog 2024-08-28 | |
| Llama 3.3 70B | FP4 trt-llm batched | 10000.0 tok/s aggregate | — | NVIDIA Blog (InferenceMAX v1 by SemiAnalysis) 2025-10-07 | |
| llama3.3:70b | NVFP4 trt-llm batched | 10000.0 tok/s aggregate | — | NVIDIA Developer Blog (SemiAnalysis InferenceMAX v1) 2025-10-09 | |
| GLM-4.5-Air | AWQ vllm batched | 9675.2 tok/s aggregate | — | Data Science Collective / Cloudrift AI 2026-01-21 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| MLPerf Inference v5.1 — Llama 2 70B Offline | 102,725 tok/s | Lambda AI / MLCommons MLPerf Inference v5.1 2025-09-09 | |
| MLPerf Inference v6.0 — gpt-oss 120B Offline | 85,921 tok/s | Nebius / MLCommons MLPerf Inference v6.0 2026-03-30 | |
| MLPerf Inference v6.0 — DeepSeek R1 Offline | 58,582 tok/s | Nebius / MLCommons MLPerf Inference v6.0 2026-03-30 |
Full Specifications
| tdp w | 1000 |
|---|---|
| segment | data-center |
| vram gb | 192 |
| vram type | HBM3e |
| cuda cores | 16896 |
| process nm | 5 |
| form factor | SXM |
| l2 cache mb | 50 |
| architecture | Blackwell |
| interconnect | NVLink 5.0 |
| memory bandwidth gbps | 8000 |
| nvlink bandwidth tbps | 1.8 |
NVIDIA B200 — Frequently Asked Questions
What is the NVIDIA B200 best used for?
When was the NVIDIA B200 released, and what was its launch MSRP?
Where do the NVIDIA B200 benchmark numbers come from?
Can the NVIDIA B200 run local LLMs?
Where can I buy the NVIDIA B200?
Buying guides that rank the NVIDIA B200's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA B200
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.
- NVIDIA H200 vs B200 for LLM Inference: What Changes in 2026
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- RTX 4090 vs RTX 5090 for Local LLM Inference: 24 GB vs 32 GB (2026)
- Prime Big Deal Days 2026 GPU Deals for Local LLMs: 12GB vs 16GB Cards
- Best GPU for gpt-oss 20B in 2026
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- Best 1440p Gaming GPUs in 2026
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- Best Retro Handhelds in 2026 — From $35 to $500
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- How to Build a Windows 98 Retro PC in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Self-hosting a Claude proxy — cache, rate-limit, and audit every request
- Budget 4K QD-Mini LED Gaming Monitors Hit New Lows in 2026
- The Adder at the Heart of Intel's 8087 Math Coprocessor
- Best Budget SSDs for Homelab and Proxmox Boot Drives in 2026
- Steam Machine 2025 Review: Couch Gaming and the 4K Question
- Best Game Controllers for Every Platform in 2026
- Open WebUI vs LM Studio: Best Local Chat Front-End for a 12GB GPU
- AI-Driven Driver Hunt: Installing Vintage Sound Cards on Windows 98 With Claude
- Best CPU Cooler for the Ryzen 7 5700X: Noctua vs DeepCool vs CoolerMaster
- Govee vs DAYBETTER vs KSIPZE: Best LED Strip for Desk and Monitor Bias Lighting
- Best Retro-PC Storage & Drive-Imaging Kit in 2026: 5 Picks
- Best $1000 Gaming PC Build for 2026: Parts We'd Actually Buy
- Period-Correct 1999 Voodoo2 SLI + Pentium III LAN-Party Build Log (2026)
- Radeon Pro W7900 48GB vs RTX A6000 48GB for Local 70B Inference (2026)
- Ryzen 7 5800X vs 5700X for Gaming and Streaming: Which AM4 Chip Wins in 2026?
- Logitech G920 vs HORI Racing Wheel: Best Starter Sim Wheel in 2026
- Cerebras Running GPT-5.4 and GPT-5.5 Internally: What the CFO's Slip Tells Us About Wafer-Scale Inference
- Steam Deck OLED vs ROG Xbox Ally X: Which Handheld Wins?
- Best Budget Gaming Mouse Pad for Esports in 2026
- Best GPU for 1440p Gaming Under $300 in 2026
- Best 1440p 240Hz Gaming Monitors for Competitive FPS in 2026
- Best Storage Upgrades for Retro and Budget PC Builds in 2026
- MAYFLASH F300 vs GameSir G7 SE: Best Controller for PC Fighting Games (2026)
- Intel Axes BigDL: Local-LLM Picks for Consumer GPUs in 2026
More buying guides from SpecPicks
Browse all buying guides →- Best Mechanical Keyboards for Gaming in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best 4K Monitors for Content Creators in 2026
- Best CPUs for Gaming in 2026
- Best CPUs for Content Creators in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best Gaming Mice for 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best PC Cases for Building in 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best External SSDs for Content Creators in 2026
- Best NVMe SSDs for Gaming in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best NVMe External Enclosures for 2026
- Best CPU Coolers for 2026
- Best Controllers for PC Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best GPUs for 4K Gaming in 2026
- Best Gaming Monitors for 2026
- Best Graphics Cards for Gaming in 2026
- Best AM5 Motherboards for 2026
Hardware benchmark data on SpecPicks
All benchmarks →- AMD Ryzen 7 7840S — benchmarks & specs
- NVIDIA GeForce RTX 5070 Ti — benchmarks & specs
- AMD Ryzen 5 7400F — benchmarks & specs
- AMD Ryzen Threadripper PRO 5955WX — benchmarks & specs
- Radeon RX 6500 XT — benchmarks & specs
- Ryzen 5 1600 — benchmarks & specs
- AMD Ryzen 7 PRO 5750G — benchmarks & specs
- AMD Ryzen Threadripper PRO 9965WX — benchmarks & specs
- Ryzen 7 9700X — benchmarks & specs
- GeForce RTX 4060 — benchmarks & specs
- NVIDIA GeForce RTX 3050 — benchmarks & specs
- AMD Ryzen 5 PRO 7640HS — benchmarks & specs
- Intel Arc A770 — benchmarks & specs
- Hailo-8 AI Processor — benchmarks & specs
- Quadro RTX 5000 — benchmarks & specs
- AMD Ryzen 7 5800U — benchmarks & specs
- Ryzen 7 9850X3D — benchmarks & specs
- Apple M3 Max 16 Core — benchmarks & specs
- AMD Ryzen 9 PRO 5945 — benchmarks & specs
- Intel Core Ultra 5 245K — benchmarks & specs
- AMD Ryzen Threadripper PRO 7945WX — benchmarks & specs
- NVIDIA GeForce RTX 2060 — benchmarks & specs
- AMD Ryzen 9 PRO 7940HS — benchmarks & specs
- GeForce RTX 5070 SUPER — benchmarks & specs