Skip to main content
NVIDIA H100 NVL 94GB
NVIDIA · GPU · Hopper

NVIDIA H100 NVL 94GB — Benchmarks & Specs

94 GB VRAM400W TDP$40,000 MSRP2023

Bottom line: how fast is the NVIDIA H100 NVL 94GB?

For local LLM inference it generates 28.2 tokens/sec running qwen2:72b at q4_K_M under ollama, per DatabaseMart. Its 94 GB of VRAM is the binding constraint for local inference: the largest model on file at 4-bit quantization on this card is mixtral:8x22b (141B parameters, q4_K_M), generating 38.3 tokens/sec under ollama, per DatabaseMart.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA H100 NVL 94GB is a datacenter accelerator from the Hopper family released in 2023 from NVIDIA. Key on-paper specs include 94 GB of HBM2e memory, 400W TDP. It launched with a $40,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 26 community AI inference reports (top 12 shown), compiled from DatabaseMart, MorphLLM, llama.cpp GitHub Discussions, Medium; each table row links to its source. Read this page when shopping the NVIDIA H100 NVL 94GB, comparing it against other datacenter accelerators for your server or cluster, or sizing it for a specific workload (LLM inference and serving, fine-tuning, or HPC compute).

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA H100 NVL 94GB, in generated tokens per second for a single request (one user's generation speed). Higher is better; each row links to the community report or benchmark database it came from. Batched serving runs are listed separately below.
Model Quantization Relative Tokens/sec VRAM used Source
gemma2:27b FP16 vllm 1574.7 tok/s 54.5 GB DatabaseMart 2025-02-01
gemma-2:27b FP16 vllm 1574.7 tok/s 54.5 GB DatabaseMart 2026-08-25
deepseek-r1-distill-qwen:32b FP16 vllm 1192.0 tok/s 65.5 GB DatabaseMart 2025-02-01
llama2:7b q4_0 llama.cpp 280.7 tok/s — llama.cpp GitHub Discussions 2024-06-01
deepseek-r1:14b q4_K_M ollama 75.0 tok/s — DatabaseMart 2025-02-01
qwen2.5-coder:7b BF16 transformers 53.1 tok/s — Medium (@wltsankalpa) 2025-09-07

Batched serving throughput

Aggregate output tokens per second on the NVIDIA H100 NVL 94GB summed across 64 concurrent requests (serving runtimes such as vLLM). This is total server throughput, not one user's generation speed — each request generates far slower — so these rows are not comparable with the single-stream table above.
Model Quantization Relative Tokens/sec (aggregate) VRAM used Source
llama3.1:8b BF16 vllm batched 12500.0 tok/s aggregate — MorphLLM 2025-01-01
deepseek-r1-distill-qwen:7b FP16 vllm batched 7314.9 tok/s aggregate — DatabaseMart 2025-03-01
deepseek-r1-distill-llama:8b FP16 vllm batched 6488.4 tok/s aggregate — DatabaseMart 2025-03-01
deepseek-r1-distill-qwen:14b FP16 vllm batched 4304.1 tok/s aggregate — DatabaseMart 2025-03-01
deepseek-r1-distill-qwen:32b FP16 vllm batched 1390.6 tok/s aggregate — DatabaseMart 2025-03-01
llama3.1:70b FP8 vllm 64 concurrent requests 460.0 tok/s aggregate — MorphLLM 2025-01-01

Full Specifications

tdp w400
vram gb94
cuda cores14592
memory typeHBM2e

NVIDIA H100 NVL 94GB — Frequently Asked Questions

What is the NVIDIA H100 NVL 94GB best used for?
NVIDIA H100 NVL 94GB is a 94 GB VRAM Hopper-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA H100 NVL 94GB released, and what was its launch MSRP?
NVIDIA H100 NVL 94GB launched in 2023 at a $40,000 MSRP. Street prices drift from launch pricing over a product's life, so check the current listing before buying.
Where do the NVIDIA H100 NVL 94GB benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA H100 NVL 94GB run local LLMs?
Yes — AI inference results for the NVIDIA H100 NVL 94GB are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 94 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA H100 NVL 94GB?
SpecPicks has no verified listing of the NVIDIA H100 NVL 94GB on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA H100 NVL 94GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA H100 NVL 94GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this datacenter accelerator.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →