NVIDIA H200 SXM 141GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA H200 SXM 141GB?
For local LLM inference it generates 59.6 tokens/sec running llama3.3:70b at FP8 under nim, per NVIDIA NIM LLMs Benchmarking. In MLPerf Inference v5.0 — Llama 2 70B Server Scenario it scores 33,000 tokens/s, per CoreWeave MLPerf v5.0 Blog. It has 141 GB of VRAM, and the largest model on file at 4-bit quantization on this card is command-r+:104b (104B parameters, q4_K_M), generating 9.8 tokens/sec under ollama, per Ollama Benchmarks. Larger models have no sourced run on file.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA H200 SXM 141GB is a datacenter accelerator from the Hopper family released in 2024 from NVIDIA. Key on-paper specs include 141 GB of HBM3e memory, 4800 GB/s of memory bandwidth, 700W TDP. Data on this page draws on 4 synthetic benchmark results, 10 community AI inference reports, 2 measured game frame-rate results, compiled from NVIDIA Technical Blog, 3DMark, NVIDIA TensorRT-LLM Blog, Ollama Benchmarks, CoreWeave MLPerf v5.0 Blog, Hardware Unboxed, HuggingFace Spaces, Millstone AI, NVIDIA NIM LLMs Benchmarking, Tom's Hardware, vllm GitHub; each table row links to its source. Read this page when shopping the NVIDIA H200 SXM 141GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).
Gaming Performance (measured FPS)
Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.
| Game | Resolution | Settings | Relative | Avg FPS | 1% low | Source |
|---|---|---|---|---|---|---|
| Call of Duty: Modern Warfare 2 | 1440p | Ultra FSR Quality | 120 fps | 95 fps | Hardware Unboxed 2023-10-22 | |
| Cyberpunk 2077 | 4K | Ultra RT on DLSS Ultra | 60 fps | 45 fps | Tom's Hardware 2023-11-15 |
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| llama3.3:70b (speculative decoding, Llama 3.2 1B draft) | FP8 TensorRT-LLM 0.15.0 | 181.7 tok/s | — | NVIDIA Technical Blog 2024-12-17 | |
| qwen3-coder:30b-a3b | FP8 vllm | 164.9 tok/s | — | Millstone AI 2026-02-03 | |
| command-r+:104b | q4_K_M ollama | 9.8 tok/s | 135.2 GB | Ollama Benchmarks 2023-10-12 |
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| llama2:13b | FP8 tensorrt-llm batched | 11819.0 tok/s aggregate | — | NVIDIA TensorRT-LLM Blog 2023-11-13 | |
| llama2:70b | FP8 tensorrt-llm batched | 3803.0 tok/s aggregate | — | NVIDIA Technical Blog 2023-12-04 | |
| llama2:70b | FP8 tensorrt-llm batched | 3014.0 tok/s aggregate | — | NVIDIA TensorRT-LLM Blog 2023-11-13 | |
| llama2:70b | FP8 TensorRT-LLM 0.5 batched | 341.0 tok/s aggregate | — | NVIDIA TensorRT-LLM GitHub (H200 launch blog) 2023-11-13 | |
| llama3.3:70b | FP8 tensorrt-llm batched | 51.1 tok/s aggregate | — | NVIDIA Technical Blog 2024-12-11 | |
| deepseek-r1:32b | q4_K_M vllm batched | 28.7 tok/s aggregate | 129.8 GB | vllm GitHub 2023-11-05 | |
| deepseek-coder-v2:236b | q4_K_M tgi batched | 14.2 tok/s aggregate | 137.9 GB | HuggingFace Spaces 2023-08-30 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| MLPerf Inference v5.0 — Llama 2 70B Server Scenario | 33,000 tokens/s | CoreWeave MLPerf v5.0 Blog 2025-04-02 | |
| MLPerf Inference v4.0 — Llama 2 70B Offline Scenario | 31,712 tokens/s | NVIDIA Developer — AI Inference Performance 2024-03-01 | |
| 3DMark Time Spy | 15,200 points | 3DMark 2023-07-15 | |
| Port Royal (RT) | 11,800 points | 3DMark 2023-07-15 |
Full Specifications
| tdp w | 700 |
|---|---|
| vram gb | 141 |
| vram type | HBM3e |
| cuda cores | 16896 |
| fp8 tflops | 3958 |
| process nm | 4 |
| form factor | SXM5 |
| nvlink gbps | 900 |
| architecture | Hopper GH100 |
| tensor cores | 528 |
| memory bandwidth gbps | 4800 |
NVIDIA H200 SXM 141GB — Frequently Asked Questions
What is the NVIDIA H200 SXM 141GB best used for?
When was the NVIDIA H200 SXM 141GB released, and what was its launch MSRP?
Where do the NVIDIA H200 SXM 141GB benchmark numbers come from?
Can the NVIDIA H200 SXM 141GB run local LLMs?
Where can I buy the NVIDIA H200 SXM 141GB?
Buying guides that rank the NVIDIA H200 SXM 141GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA H200 SXM 141GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.
- NVIDIA H200 vs B200 for LLM Inference: What Changes in 2026
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Prime Big Deal Days 2026 GPU Deals for Local LLMs: 12GB vs 16GB Cards
- Ryzen 5 5600X vs Ryzen 5 5600G for CPU-Only Gemma 3 12B
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- RTX 4090 vs RTX 5090 for Local LLM Inference: 24 GB vs 32 GB (2026)
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- Best 1440p Gaming GPUs in 2026
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- How to Build a Windows 98 Retro PC in 2026
- Best Retro Handhelds in 2026 — From $35 to $500
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
More reviews from the SpecPicks archive
Browse all reviews →- Local LLM vs Claude in 2026: What an RTX 3060 12GB Rig Actually Replaces
- Best Budget Gaming Mouse Pad for Esports in 2026
- Best Retro Gaming Gifts in 2026: 5 Picks for Console and PC Collectors
- Best Budget GPU for CNN & Vision Inference 2026: RTX 3060 12GB
- Steam Deck OLED vs ROG Ally X vs Legion Go: 2026 Gaming Handheld Shootout
- Intel-Scaler vLLM 0.21.0: What the New Release Changes for Intel GPU Inference
- Raspberry Pi 500 Specs, Price & Real-World Use in 2026
- Forza Horizon 6 Boots in 4 Seconds with Advanced Shader Delivery
- Ryzen 5 5600G vs Ryzen 7 5700X vs Ryzen 7 5800X for 1080p Gaming in 2026
- Forza Horizon 6 Advanced Shader Delivery: 4-Second Loads vs 90 Seconds Explained
- Best GPU for Llama 3.1 405B (2026)
- How to run Qwen 3 32B on AMD Radeon RX 7900 XTX
- 20GB Homelab Memory Left: What to Add Next
- Inkling 975B: Thinking Machines' Open-Weight Speech Model, Explained
- Two GTX 1050 Ti 4GB vs One RTX 3060 12GB for 8B Local LLMs
- Logitech Gaming Mouse Deals Hit 47% Off for Prime Day 2026
- Best PC Couch-Gaming Setup in 2026: Controllers, Display, and Audio
- AMD Ryzen AI Halo Ships with a Fully Open-Source Linux Stack
- Imaging Vintage Hard Drives in 2026: SATA/IDE-to-USB Adapters Compared (FIDECO vs Unitek vs Vantec)
- Aider vs Cline vs Cursor for AI-Assisted Coding in 2026
- GeForce FX 5900 Ultra vs Radeon 9800 Pro: The 2003 DirectX 9 Showdown Revisited
- Intel Axes BigDL/IPEX-LLM: Where Local Inference Goes Now
- Is the KOORUI 27" 4K QD-Mini LED Worth It for RTX 3060 Gaming?
- Athlon XP + Radeon 9800 Pro AGP Build Guide: 2003 Period-Correct Rig
More buying guides from SpecPicks
Browse all buying guides →- Best CPUs for Gaming in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best Gaming Mice for 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best Gaming Monitors for 2026
- Best 4K Monitors for Content Creators in 2026
- Best CPU Coolers for 2026
- Best Graphics Cards for Gaming in 2026
- Best NVMe External Enclosures for 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best GPUs for 4K Gaming in 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best NVMe SSDs for Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best AM5 Motherboards for 2026
- Best PC Cases for Building in 2026
- Best Mechanical Keyboards for Gaming in 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best Controllers for PC Gaming in 2026
- Best External SSDs for Content Creators in 2026
- Best CPUs for Content Creators in 2026
Hardware benchmark data on SpecPicks
All benchmarks →- NVIDIA L40S 48GB — benchmarks & specs
- Ryzen 5 5500X3D — benchmarks & specs
- NVIDIA B200 — benchmarks & specs
- Ryzen 5 7500F — benchmarks & specs
- AMD Ryzen 5 5600U — benchmarks & specs
- GRID RTX6000-6Q — benchmarks & specs
- Ryzen 7 3700X — benchmarks & specs
- AMD Ryzen 5 7530U — benchmarks & specs
- AMD Ryzen 7 PRO 7745 — benchmarks & specs
- AMD Ryzen 7 5825C — benchmarks & specs
- Apple M3 Ultra 32 Core — benchmarks & specs
- Radeon RX 6800S — benchmarks & specs
- Radeon RX 6700M — benchmarks & specs
- GeForce RTX 5060 Ti 16GB — benchmarks & specs
- AMD Ryzen 5 7545U — benchmarks & specs
- Radeon RX 6600S — benchmarks & specs
- Apple M2 Pro 10 Core 3480 MHz — benchmarks & specs
- AMD Ryzen Threadripper PRO 5995WX — benchmarks & specs
- NVIDIA Jetson Orin Nano Super — benchmarks & specs
- AMD Ryzen 3 5300GE — benchmarks & specs
- Ryzen 7 4800H — benchmarks & specs
- GRID RTX6000-24Q — benchmarks & specs
- NVIDIA H100 PCIe 80GB — benchmarks & specs
- RadeonT RX 6850M XT — benchmarks & specs