Skip to main content
NVIDIA H100 PCIe 80GB
NVIDIA · GPU · Hopper

NVIDIA H100 PCIe 80GB — Benchmarks & Specs

80 GB VRAM350W TDP$32,000 MSRP2022

Bottom line: how fast is the NVIDIA H100 PCIe 80GB?

For local LLM inference it generates 28.2 tokens/sec running qwen3:72b at q4_K_M under ollama, per DatabaseMart. In 3DMark Time Spy it scores 2,681 points, per TechSpot. Its 80 GB of VRAM is the binding constraint for local inference: the largest model on file at 4-bit quantization on this card is qwen2.5:110b (110B parameters, q4_K_M), generating 20.2 tokens/sec under ollama, per DatabaseMart.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA H100 PCIe 80GB is a datacenter accelerator from the Hopper family released in 2022 from NVIDIA. Key on-paper specs include 80 GB of HBM2e memory, 350W TDP. It launched with a $32,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 1 synthetic benchmark result, 26 community AI inference reports (top 12 shown), 1 measured game frame-rate result, compiled from DatabaseMart, TechSpot, GPU-Benchmarks-on-LLM-Inference, Koyeb GPU Benchmarks, llama.cpp GitHub Discussions, Spheron Blog, VALDI Docs; each table row links to its source. Read this page when shopping the NVIDIA H100 PCIe 80GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).

Gaming Performance (measured FPS)

Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.

Measured gaming frame rates for the NVIDIA H100 PCIe 80GB by game, resolution, and quality preset. “1% low” is the frame-time floor that determines perceived smoothness. Each row links to its original review or benchmark database.
Game Resolution Settings Relative Avg FPS 1% low Source
Red Dead Redemption 2 1440p High 8 fps — TechSpot 2023-06-21

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA H100 PCIe 80GB, in generated tokens per second for a single request (one user's generation speed). Higher is better; each row links to the community report or benchmark database it came from. Batched serving runs are listed separately below.
Model Quantization Relative Tokens/sec VRAM used Source
deepseek-r1-distill-qwen:14b FP16 vllm 3689.2 tok/s — DatabaseMart 2026-08-25
gemma2:27b BF16 vllm 1575.0 tok/s 54.5 GB DatabaseMart 2025-03-01
llama2:7b q4_0 llama.cpp 280.7 tok/s — llama.cpp GitHub Discussions 2024-01-01
llama3:8b Q4_K_M llama.cpp 145.6 tok/s — XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-12-01
llama3:8b q4_K_M llama.cpp 144.5 tok/s — GPU-Benchmarks-on-LLM-Inference (GitHub, XiongjieDai) 2024-05-01
llama3.3:70b FP8 vllm 120.0 tok/s 76.0 GB Spheron Blog 2026-01-01

Batched serving throughput

Aggregate output tokens per second on the NVIDIA H100 PCIe 80GB summed across many concurrent requests (serving runtimes such as vLLM). This is total server throughput, not one user's generation speed — each request generates far slower — so these rows are not comparable with the single-stream table above.
Model Quantization Relative Tokens/sec (aggregate) VRAM used Source
deepseek-r1-distill-qwen:7b FP16 vllm batched 6270.0 tok/s aggregate 72.0 GB DatabaseMart 2025-03-01
deepseek-r1-distill-llama:8b FP16 vllm batched 5561.0 tok/s aggregate 72.0 GB DatabaseMart 2025-03-01
deepseek-r1-distill-qwen:14b FP16 vllm batched 3689.0 tok/s aggregate 72.0 GB DatabaseMart 2025-03-01
llama3.1:8b BF16 vllm 64 concurrent requests 3621.0 tok/s aggregate — VALDI Docs 2024-09-01
llama3.1:8b FP16 vllm 32 concurrent requests 3008.0 tok/s aggregate — Koyeb GPU Benchmarks 2025-01-01
deepseek-r1-distill-qwen:32b FP16 vllm batched 1192.0 tok/s aggregate 72.0 GB DatabaseMart 2025-03-01

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA H100 PCIe 80GB — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
3DMark Time Spy 2,681 points TechSpot 2023-06-21

Full Specifications

tdp w350
vram gb80
cuda cores14592
memory typeHBM2e

NVIDIA H100 PCIe 80GB — Frequently Asked Questions

What is the NVIDIA H100 PCIe 80GB best used for?
NVIDIA H100 PCIe 80GB is a 80 GB VRAM Hopper-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA H100 PCIe 80GB released, and what was its launch MSRP?
NVIDIA H100 PCIe 80GB launched in 2022 at a $32,000 MSRP. Street prices drift from launch pricing over a product's life, so check the current listing before buying.
Where do the NVIDIA H100 PCIe 80GB benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA H100 PCIe 80GB run local LLMs?
Yes — AI inference results for the NVIDIA H100 PCIe 80GB are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 80 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA H100 PCIe 80GB?
SpecPicks has no verified listing of the NVIDIA H100 PCIe 80GB on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA H100 PCIe 80GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA H100 PCIe 80GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →