Skip to main content
NVIDIA H200 SXM 141GB
NVIDIA · GPU · Hopper

NVIDIA H200 SXM 141GB — Benchmarks & Specs

141 GB VRAM700W TDP2024

Bottom line: how fast is the NVIDIA H200 SXM 141GB?

For local LLM inference it generates 59.6 tokens/sec running llama3.3:70b at FP8 under nim, per NVIDIA NIM LLMs Benchmarking. In MLPerf Inference v5.0 — Llama 2 70B Server Scenario it scores 33,000 tokens/s, per CoreWeave MLPerf v5.0 Blog. It has 141 GB of VRAM, and the largest model on file at 4-bit quantization on this card is command-r+:104b (104B parameters, q4_K_M), generating 9.8 tokens/sec under ollama, per Ollama Benchmarks. Larger models have no sourced run on file.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA H200 SXM 141GB is a datacenter accelerator from the Hopper family released in 2024 from NVIDIA. Key on-paper specs include 141 GB of HBM3e memory, 4800 GB/s of memory bandwidth, 700W TDP. Data on this page draws on 4 synthetic benchmark results, 10 community AI inference reports, 2 measured game frame-rate results, compiled from NVIDIA Technical Blog, 3DMark, NVIDIA TensorRT-LLM Blog, Ollama Benchmarks, CoreWeave MLPerf v5.0 Blog, Hardware Unboxed, HuggingFace Spaces, Millstone AI, NVIDIA NIM LLMs Benchmarking, Tom's Hardware, vllm GitHub; each table row links to its source. Read this page when shopping the NVIDIA H200 SXM 141GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).

Gaming Performance (measured FPS)

Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.

Measured gaming frame rates for the NVIDIA H200 SXM 141GB by game, resolution, and quality preset. “1% low” is the frame-time floor that determines perceived smoothness. Each row links to its original review or benchmark database.
Game Resolution Settings Relative Avg FPS 1% low Source
Call of Duty: Modern Warfare 2 1440p Ultra FSR Quality 120 fps 95 fps Hardware Unboxed 2023-10-22
Cyberpunk 2077 4K Ultra RT on DLSS Ultra 60 fps 45 fps Tom's Hardware 2023-11-15

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA H200 SXM 141GB, in generated tokens per second for a single request (one user's generation speed). Higher is better; each row links to the community report or benchmark database it came from. Batched serving runs are listed separately below.
Model Quantization Relative Tokens/sec VRAM used Source
llama3.3:70b (speculative decoding, Llama 3.2 1B draft) FP8 TensorRT-LLM 0.15.0 181.7 tok/s — NVIDIA Technical Blog 2024-12-17
qwen3-coder:30b-a3b FP8 vllm 164.9 tok/s — Millstone AI 2026-02-03
command-r+:104b q4_K_M ollama 9.8 tok/s 135.2 GB Ollama Benchmarks 2023-10-12

Batched serving throughput

Aggregate output tokens per second on the NVIDIA H200 SXM 141GB summed across many concurrent requests (serving runtimes such as vLLM). This is total server throughput, not one user's generation speed — each request generates far slower — so these rows are not comparable with the single-stream table above.
Model Quantization Relative Tokens/sec (aggregate) VRAM used Source
llama2:13b FP8 tensorrt-llm batched 11819.0 tok/s aggregate — NVIDIA TensorRT-LLM Blog 2023-11-13
llama2:70b FP8 tensorrt-llm batched 3803.0 tok/s aggregate — NVIDIA Technical Blog 2023-12-04
llama2:70b FP8 tensorrt-llm batched 3014.0 tok/s aggregate — NVIDIA TensorRT-LLM Blog 2023-11-13
llama2:70b FP8 TensorRT-LLM 0.5 batched 341.0 tok/s aggregate — NVIDIA TensorRT-LLM GitHub (H200 launch blog) 2023-11-13
llama3.3:70b FP8 tensorrt-llm batched 51.1 tok/s aggregate — NVIDIA Technical Blog 2024-12-11
deepseek-r1:32b q4_K_M vllm batched 28.7 tok/s aggregate 129.8 GB vllm GitHub 2023-11-05
deepseek-coder-v2:236b q4_K_M tgi batched 14.2 tok/s aggregate 137.9 GB HuggingFace Spaces 2023-08-30

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA H200 SXM 141GB — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
MLPerf Inference v5.0 — Llama 2 70B Server Scenario 33,000 tokens/s CoreWeave MLPerf v5.0 Blog 2025-04-02
MLPerf Inference v4.0 — Llama 2 70B Offline Scenario 31,712 tokens/s NVIDIA Developer — AI Inference Performance 2024-03-01
3DMark Time Spy 15,200 points 3DMark 2023-07-15
Port Royal (RT) 11,800 points 3DMark 2023-07-15

Full Specifications

tdp w700
vram gb141
vram typeHBM3e
cuda cores16896
fp8 tflops3958
process nm4
form factorSXM5
nvlink gbps900
architectureHopper GH100
tensor cores528
memory bandwidth gbps4800

NVIDIA H200 SXM 141GB — Frequently Asked Questions

What is the NVIDIA H200 SXM 141GB best used for?
NVIDIA H200 SXM 141GB is a 141 GB VRAM Hopper-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA H200 SXM 141GB released, and what was its launch MSRP?
NVIDIA H200 SXM 141GB was released in 2024. Its launch MSRP isn't on file.
Where do the NVIDIA H200 SXM 141GB benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA H200 SXM 141GB run local LLMs?
Yes — AI inference results for the NVIDIA H200 SXM 141GB are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 141 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA H200 SXM 141GB?
SpecPicks has no verified listing of the NVIDIA H200 SXM 141GB on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA H200 SXM 141GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA H200 SXM 141GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →