Skip to main content
NVIDIA B200
NVIDIA · GPU · Blackwell

NVIDIA B200 — Benchmarks & Specs

192 GB VRAM1000W TDP2024

Bottom line: how fast is the NVIDIA B200?

In a batched serving run it reaches 12841.0 tokens/sec aggregate throughput (many concurrent requests, not one user's generation speed) running llama2:70b at FP4 under trt-llm, per Lambda AI. In MLPerf Inference v5.1 — Llama 2 70B Offline it scores 102,725 tok/s, per Lambda AI. It has 192 GB of VRAM, and the largest model on file at 4-bit quantization on this card is gpt-oss-120b (117B parameters, MXFP4), generating 60000.0 tokens/sec under trt-llm, per NVIDIA Developer Blog. Larger models have no sourced run on file.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA B200 is a datacenter accelerator from the Blackwell family released in 2024 from NVIDIA. Key on-paper specs include 192 GB of HBM3e memory, 8000 GB/s of memory bandwidth, 1000W TDP. Data on this page draws on 3 synthetic benchmark results, 6 community AI inference reports, compiled from NVIDIA Developer Blog, Nebius / MLCommons MLPerf Inference v6.0, Data Science Collective / Cloudrift AI, Lambda AI, NVIDIA Blog; each table row links to its source. Read this page when shopping the NVIDIA B200, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Batched serving throughput

Aggregate output tokens per second on the NVIDIA B200 summed across many concurrent requests (serving runtimes such as vLLM). This is total server throughput, not one user's generation speed — each request generates far slower — so these rows are not comparable with the single-stream table.
Model Quantization Relative Tokens/sec (aggregate) VRAM used Source
gpt-oss-120b MXFP4 trt-llm batched 60000.0 tok/s aggregate — NVIDIA Developer Blog (SemiAnalysis InferenceMAX v1) 2025-10-09
gpt-oss-120b NVFP4 tgi batched 30000.0 tok/s aggregate — NVIDIA Developer Blog / SemiAnalysis InferenceMAX 2025-10-13
llama2:70b FP4 trt-llm batched 11264.0 tok/s aggregate — NVIDIA Developer Blog 2024-08-28
Llama 3.3 70B FP4 trt-llm batched 10000.0 tok/s aggregate — NVIDIA Blog (InferenceMAX v1 by SemiAnalysis) 2025-10-07
llama3.3:70b NVFP4 trt-llm batched 10000.0 tok/s aggregate — NVIDIA Developer Blog (SemiAnalysis InferenceMAX v1) 2025-10-09
GLM-4.5-Air AWQ vllm batched 9675.2 tok/s aggregate — Data Science Collective / Cloudrift AI 2026-01-21

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA B200 — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
MLPerf Inference v5.1 — Llama 2 70B Offline 102,725 tok/s Lambda AI / MLCommons MLPerf Inference v5.1 2025-09-09
MLPerf Inference v6.0 — gpt-oss 120B Offline 85,921 tok/s Nebius / MLCommons MLPerf Inference v6.0 2026-03-30
MLPerf Inference v6.0 — DeepSeek R1 Offline 58,582 tok/s Nebius / MLCommons MLPerf Inference v6.0 2026-03-30

Full Specifications

tdp w1000
segmentdata-center
vram gb192
vram typeHBM3e
cuda cores16896
process nm5
form factorSXM
l2 cache mb50
architectureBlackwell
interconnectNVLink 5.0
memory bandwidth gbps8000
nvlink bandwidth tbps1.8

NVIDIA B200 — Frequently Asked Questions

What is the NVIDIA B200 best used for?
NVIDIA B200 is a 192 GB VRAM Blackwell-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA B200 released, and what was its launch MSRP?
NVIDIA B200 was released in 2024. Its launch MSRP isn't on file.
Where do the NVIDIA B200 benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA B200 run local LLMs?
Yes — AI inference results for the NVIDIA B200 are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 192 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA B200?
SpecPicks has no verified listing of the NVIDIA B200 on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA B200's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA B200

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →