NVIDIA H100 PCIe 80GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA H100 PCIe 80GB?
For local LLM inference it generates 28.2 tokens/sec running qwen3:72b at q4_K_M under ollama, per DatabaseMart. In 3DMark Time Spy it scores 2,681 points, per TechSpot. Its 80 GB of VRAM is the binding constraint for local inference: the largest model on file at 4-bit quantization on this card is qwen2.5:110b (110B parameters, q4_K_M), generating 20.2 tokens/sec under ollama, per DatabaseMart.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA H100 PCIe 80GB is a datacenter accelerator from the Hopper family released in 2022 from NVIDIA. Key on-paper specs include 80 GB of HBM2e memory, 350W TDP. It launched with a $32,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 1 synthetic benchmark result, 26 community AI inference reports (top 12 shown), 1 measured game frame-rate result, compiled from DatabaseMart, TechSpot, GPU-Benchmarks-on-LLM-Inference, Koyeb GPU Benchmarks, llama.cpp GitHub Discussions, Spheron Blog, VALDI Docs; each table row links to its source. Read this page when shopping the NVIDIA H100 PCIe 80GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).
Gaming Performance (measured FPS)
Average and 1% low frame rates by game, resolution, and quality preset. Bars are scaled against the fastest result on this page.
| Game | Resolution | Settings | Relative | Avg FPS | 1% low | Source |
|---|---|---|---|---|---|---|
| Red Dead Redemption 2 | 1440p | High | 8 fps | — | TechSpot 2023-06-21 |
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| deepseek-r1-distill-qwen:14b | FP16 vllm | 3689.2 tok/s | — | DatabaseMart 2026-08-25 | |
| gemma2:27b | BF16 vllm | 1575.0 tok/s | 54.5 GB | DatabaseMart 2025-03-01 | |
| llama2:7b | q4_0 llama.cpp | 280.7 tok/s | — | llama.cpp GitHub Discussions 2024-01-01 | |
| llama3:8b | Q4_K_M llama.cpp | 145.6 tok/s | — | XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-12-01 | |
| llama3:8b | q4_K_M llama.cpp | 144.5 tok/s | — | GPU-Benchmarks-on-LLM-Inference (GitHub, XiongjieDai) 2024-05-01 | |
| llama3.3:70b | FP8 vllm | 120.0 tok/s | 76.0 GB | Spheron Blog 2026-01-01 |
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| deepseek-r1-distill-qwen:7b | FP16 vllm batched | 6270.0 tok/s aggregate | 72.0 GB | DatabaseMart 2025-03-01 | |
| deepseek-r1-distill-llama:8b | FP16 vllm batched | 5561.0 tok/s aggregate | 72.0 GB | DatabaseMart 2025-03-01 | |
| deepseek-r1-distill-qwen:14b | FP16 vllm batched | 3689.0 tok/s aggregate | 72.0 GB | DatabaseMart 2025-03-01 | |
| llama3.1:8b | BF16 vllm 64 concurrent requests | 3621.0 tok/s aggregate | — | VALDI Docs 2024-09-01 | |
| llama3.1:8b | FP16 vllm 32 concurrent requests | 3008.0 tok/s aggregate | — | Koyeb GPU Benchmarks 2025-01-01 | |
| deepseek-r1-distill-qwen:32b | FP16 vllm batched | 1192.0 tok/s aggregate | 72.0 GB | DatabaseMart 2025-03-01 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| 3DMark Time Spy | 2,681 points | TechSpot 2023-06-21 |
Full Specifications
| tdp w | 350 |
|---|---|
| vram gb | 80 |
| cuda cores | 14592 |
| memory type | HBM2e |
NVIDIA H100 PCIe 80GB — Frequently Asked Questions
What is the NVIDIA H100 PCIe 80GB best used for?
When was the NVIDIA H100 PCIe 80GB released, and what was its launch MSRP?
Where do the NVIDIA H100 PCIe 80GB benchmark numbers come from?
Can the NVIDIA H100 PCIe 80GB run local LLMs?
Where can I buy the NVIDIA H100 PCIe 80GB?
Buying guides that rank the NVIDIA H100 PCIe 80GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA H100 PCIe 80GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Ryzen 5 5600X vs Ryzen 5 5600G for CPU-Only Gemma 3 12B
- Prime Big Deal Days 2026 GPU Deals for Local LLMs: 12GB vs 16GB Cards
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Best Hardware for Local OCR and Document AI in 2026
- RTX 4090 vs RTX 5090 for Local LLM Inference: 24 GB vs 32 GB (2026)
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- Best 1440p Gaming GPUs in 2026
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- Best Retro Handhelds in 2026 — From $35 to $500
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- How to Build a Windows 98 Retro PC in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Best GPU for 1440p Gaming Under $300 in 2026
- Best Storage for a Raspberry Pi 4 Home Server: microSD vs SATA SSD vs NVMe
- Microsoft + Nvidia Agent PCs: Hardware to Run Agents Locally
- Ryzen 7 5800X vs i7-9700K: Better Home Workstation in 2026?
- RetroPie on a Raspberry Pi 4 8GB: The Definitive 2026 Emulation Build
- When a Game Boy Clone Runs Too Fast: The Crystal-Swap Fix Behind a Viral Repair
- Best Handheld Gaming PC Under $300 in 2026
- How to run Llama 3.1 8B on AMD Radeon RX 7900 XTX
- HiDream-O1-Image on an RTX 3060 12GB: Does It Fit?
- Raspberry Pi 5 vs Pi 4 8GB for Homelab: Which Should You Buy?
- VibeThinker-3B Local: 3B Reasoning Model on an RTX 3060 12GB
- Leanstral 1.5 on an RTX 3060 12GB: Local Math + Bug-Finding Benchmarks
- Steam Deck OLED vs ROG Ally X vs Legion Go: 2026 Gaming Handheld Shootout
- RTX 3060 12GB vs RTX 3090 for Local LLMs (2026)
- Kimi K2.7 Code Is 12x Cheaper Than GPT-5.5 — Run It Local?
- oQ vs Q4_K/Q6_K vs MXFP4 vs UD-MLX: Which Quantization Format Should You Pick in 2026?
- Raspberry Pi Zero 2W: Privacy-Preserving Ring Alternative
- Imaging Vintage IDE Drives with a CompactFlash + USB Adapter
- Open-Weight Models Caught Up to Frontier: What to Run on a 12GB GPU
- Raspberry Pi 5 Specs: Full 2025 Breakdown
- Build a Live ADS-B Flight Tracker on a Raspberry Pi 4 in 2026
- 3dfx Voodoo2 SLI on Windows 98 SE: Period-Correct Build, Glide Driver Setup, and Real Quake 3 Benchmarks
- Gemma 4 Stealth Update Fixes Tool Calling: What Changes Locally
- Self-Hosting Jellyfin on a Raspberry Pi 4 8GB: Real Performance
More buying guides from SpecPicks
Browse all buying guides →- Best Gaming Monitors for 2026
- Best External SSDs for Content Creators in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best AM5 Motherboards for 2026
- Best GPUs for 4K Gaming in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best CPU Coolers for 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best Gaming Mice for 2026
- Best NVMe SSDs for Gaming in 2026
- Best 4K Monitors for Content Creators in 2026
- Best PC Cases for Building in 2026
- Best CPUs for Gaming in 2026
- Best Mechanical Keyboards for Gaming in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best CPUs for Content Creators in 2026
- Best NVMe External Enclosures for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best Controllers for PC Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best Graphics Cards for Gaming in 2026
Hardware benchmark data on SpecPicks
All benchmarks →- Apple M2 Max 12 Core 3680 MHz — benchmarks & specs
- GRID RTX6000-1B — benchmarks & specs
- Apple M3 Max 16 Core — benchmarks & specs
- GRID RTX6000P-2B — benchmarks & specs
- DDR5-6400 CL32 32GB (2x16) — benchmarks & specs
- GeForce RTX 5050 9 GB — benchmarks & specs
- AMD Ryzen 3 7335U — benchmarks & specs
- AMD Ryzen 7 7435HS — benchmarks & specs
- NVIDIA GeForce RTX 3090 — benchmarks & specs
- NVIDIA GeForce GTX 1660 SUPER — benchmarks & specs
- AMD Ryzen Threadripper PRO 7965WX — benchmarks & specs
- NVIDIA A100 PCIe 80GB — benchmarks & specs
- Intel Arc A770M — benchmarks & specs
- Ryzen 3 3100 — benchmarks & specs
- AMD Ryzen 9 9955HX3D — benchmarks & specs
- GeForce RTX 5060 Ti 16GB — benchmarks & specs
- Apple M4 Pro 12 Core — benchmarks & specs
- NVIDIA Jetson Orin Nano 8GB — benchmarks & specs
- NVIDIA GeForce RTX 4070 Ti SUPER — benchmarks & specs
- Ryzen 5 3600X — benchmarks & specs
- Radeon RX 7900M — benchmarks & specs
- RTX 6000D — benchmarks & specs
- RTX5000-Ada-4Q — benchmarks & specs
- AMD Ryzen 3 PRO 5355G — benchmarks & specs