NVIDIA H100 NVL 94GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA H100 NVL 94GB?
For local LLM inference it generates 28.2 tokens/sec running qwen2:72b at q4_K_M under ollama, per DatabaseMart. Its 94 GB of VRAM is the binding constraint for local inference: the largest model on file at 4-bit quantization on this card is mixtral:8x22b (141B parameters, q4_K_M), generating 38.3 tokens/sec under ollama, per DatabaseMart.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA H100 NVL 94GB is a datacenter accelerator from the Hopper family released in 2023 from NVIDIA. Key on-paper specs include 94 GB of HBM2e memory, 400W TDP. It launched with a $40,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 26 community AI inference reports (top 12 shown), compiled from DatabaseMart, MorphLLM, llama.cpp GitHub Discussions, Medium; each table row links to its source. Read this page when shopping the NVIDIA H100 NVL 94GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| gemma2:27b | FP16 vllm | 1574.7 tok/s | 54.5 GB | DatabaseMart 2025-02-01 | |
| gemma-2:27b | FP16 vllm | 1574.7 tok/s | 54.5 GB | DatabaseMart 2026-08-25 | |
| deepseek-r1-distill-qwen:32b | FP16 vllm | 1192.0 tok/s | 65.5 GB | DatabaseMart 2025-02-01 | |
| llama2:7b | q4_0 llama.cpp | 280.7 tok/s | — | llama.cpp GitHub Discussions 2024-06-01 | |
| deepseek-r1:14b | q4_K_M ollama | 75.0 tok/s | — | DatabaseMart 2025-02-01 | |
| qwen2.5-coder:7b | BF16 transformers | 53.1 tok/s | — | Medium (@wltsankalpa) 2025-09-07 |
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| llama3.1:8b | BF16 vllm batched | 12500.0 tok/s aggregate | — | MorphLLM 2025-01-01 | |
| deepseek-r1-distill-qwen:7b | FP16 vllm batched | 7314.9 tok/s aggregate | — | DatabaseMart 2025-03-01 | |
| deepseek-r1-distill-llama:8b | FP16 vllm batched | 6488.4 tok/s aggregate | — | DatabaseMart 2025-03-01 | |
| deepseek-r1-distill-qwen:14b | FP16 vllm batched | 4304.1 tok/s aggregate | — | DatabaseMart 2025-03-01 | |
| deepseek-r1-distill-qwen:32b | FP16 vllm batched | 1390.6 tok/s aggregate | — | DatabaseMart 2025-03-01 | |
| llama3.1:70b | FP8 vllm 64 concurrent requests | 460.0 tok/s aggregate | — | MorphLLM 2025-01-01 |
Full Specifications
| tdp w | 400 |
|---|---|
| vram gb | 94 |
| cuda cores | 14592 |
| memory type | HBM2e |
NVIDIA H100 NVL 94GB — Frequently Asked Questions
What is the NVIDIA H100 NVL 94GB best used for?
When was the NVIDIA H100 NVL 94GB released, and what was its launch MSRP?
Where do the NVIDIA H100 NVL 94GB benchmark numbers come from?
Can the NVIDIA H100 NVL 94GB run local LLMs?
Where can I buy the NVIDIA H100 NVL 94GB?
Buying guides that rank the NVIDIA H100 NVL 94GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA H100 NVL 94GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Best GPU for gpt-oss 20B in 2026
- Ryzen 5 5600X vs Ryzen 5 5600G for CPU-Only Gemma 3 12B
- Best Hardware for Local OCR and Document AI in 2026
- RTX 4090 vs RTX 5090 for Local LLM Inference: 24 GB vs 32 GB (2026)
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- Best Retro Handhelds in 2026 — From $35 to $500
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- How to Build a Windows 98 Retro PC in 2026
- Best 1440p Gaming GPUs in 2026
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Jetson Orin Nano Super vs Raspberry Pi 5: Real Edge-AI Benchmarks (2026)
- Ryzen AI Max Mini PC 128GB: Specs, Uses, Alternatives
- 8BitDo Pro 2 vs GameSir G7 SE for Raspberry Pi Emulation
- AI-Assisted Driver Hunting on Voodoo3 + GeForce 4 Ti: A 2026 Win98 Workflow
- Best PC Streaming Starter Kit in 2026: 5 Picks
- RTX 5080 vs RTX 4080 for LLM Inference: Same 16GB, Different Answer (2026)
- Best Retro PC Upgrade Kit 2026: Sound & Storage
- Best Internal SSD for a PS4 Pro Upgrade: 2.5" SATA Picks That Cut Load Times
- Best SSD for Retro PC IDE-to-SATA Builds (Win98, WinXP) in 2026
- Silent SSD-Style Windows 98 Build: CompactFlash Boot + IDE Adapter Workflow
- Best Budget PC Gaming Peripherals in 2026: 5 Picks That Punch Above Their Price
- Build a Pocket Retro Emulation Handheld on the Raspberry Pi Zero W in 2026
- Intel Arc LLM Inference: A770 & B580 Guide 2025
- Best Budget Gaming Monitors Under $300 for 1080p in 2026
- Self-Hosted Jellyfin on a Raspberry Pi 4 8GB: What 1080p Transcoding Really Costs
- Best SSD for a Big Steam Library: Crucial BX500 1TB vs WD Blue
- Best PC Gaming Controller in 2026: GameSir G7 SE vs DualSense vs 8BitDo Pro 2
- FIDECO vs Unitek vs Vantec: Best IDE/SATA-to-USB Adapter for Retro PCs
- Best SSD for Raspberry Pi Boot (2026): USB SATA vs NVMe
- Best Gaming Headset for PS5 and PC in 2026
- Self-Host Home Assistant on a Raspberry Pi 4 8GB: Setup and Real Limits
- First-Time CPU Delidding: What to Expect and How to Get It Right
- Etched's Transformer-Only Inference Chip vs Your GPU: What Changes for Local Builders
- Best Buy Cuts $900 Off an Asus 64GB Gaming Laptop — Is a Desktop Smarter?
More buying guides from SpecPicks
Browse all buying guides →- Best NVMe SSDs for Gaming in 2026
- Best CPUs for Gaming in 2026
- Best AM5 Motherboards for 2026
- Best GPUs for 4K Gaming in 2026
- Best Gaming Monitors for 2026
- Best CPUs for Content Creators in 2026
- Best 4K Monitors for Content Creators in 2026
- Best CPU Coolers for 2026
- Best Controllers for PC Gaming in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best Graphics Cards for Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best External SSDs for Content Creators in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best PC Cases for Building in 2026
- Best Gaming Mice for 2026
- Best NVMe External Enclosures for 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best Mechanical Keyboards for Gaming in 2026
- Best DDR5 RAM for Gaming PCs in 2026
Hardware benchmark data on SpecPicks
All benchmarks →- GRID RTX6000-8Q — benchmarks & specs
- AMD Ryzen Threadripper PRO 7995WX — benchmarks & specs
- AMD Ryzen 3 5300U — benchmarks & specs
- Radeon RX 6800M — benchmarks & specs
- Ryzen 5 5500X3D — benchmarks & specs
- Ryzen 5 3500X — benchmarks & specs
- Apple M2 8 Core 3500 MHz — benchmarks & specs
- NVIDIA GeForce RTX 5090 — benchmarks & specs
- AMD Ryzen 7 PRO 7840HS — benchmarks & specs
- AMD Ryzen 9 PRO 9945 — benchmarks & specs
- AMD Ryzen 5 7530U — benchmarks & specs
- AMD Ryzen Threadripper PRO 7965WX — benchmarks & specs
- Apple M1 Pro 8 Core 3200 MHz — benchmarks & specs
- Intel Arc A530M — benchmarks & specs
- Intel Arc A370M — benchmarks & specs
- Radeon RX 7600 — benchmarks & specs
- Radeon RX 7900M — benchmarks & specs
- NVIDIA GeForce RTX 4090 — benchmarks & specs
- AMD Ryzen 7 7840H — benchmarks & specs
- NVIDIA GeForce RTX 3090 Ti — benchmarks & specs
- AMD Ryzen 7 7435H — benchmarks & specs
- AMD Ryzen 7 PRO 5755G — benchmarks & specs
- Ryzen 3 2200G — benchmarks & specs
- Radeon RX 7900 XT — benchmarks & specs