Skip to main content
NVIDIA A100 PCIe 40GB
NVIDIA · GPU · Ampere

NVIDIA A100 PCIe 40GB — Benchmarks & Specs

40 GB VRAM250W TDP$12,000 MSRP2020

Bottom line: how fast is the NVIDIA A100 PCIe 40GB?

For local LLM inference it generates 3385.7 tokens/sec running gemma-3-4b-it at FP16 under vllm, per DatabaseMart. In PassMark G3D Mark it scores 13,326 points, per PassMark VideoCardBenchmark. Its 40 GB of VRAM is the binding constraint for local inference: that capacity fits 32B-parameter models at Q4 without offloading to system RAM.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA A100 PCIe 40GB is a graphics card from the Ampere family released in 2020 from NVIDIA. Key on-paper specs include 40 GB of GDDR VRAM, 250W TDP. It launched with a $12,000 MSRP, though street prices typically diverge meaningfully from launch pricing — see the linked product cards below for current Amazon listings. Data on this page draws on 5 synthetic benchmark results, 12 community AI inference reports, aggregated from public benchmark databases (TechPowerUp, PassMark, Geekbench, Cinebench) and the LocalLLaMA community. Read this page when shopping the NVIDIA A100 PCIe 40GB, comparing it against other graphics cards in your build, or sizing it for a specific workload (gaming at 1080p/1440p/4K, productivity benchmarks, or local LLM inference).

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA A100 PCIe 40GB, in generated tokens per second. Higher is better; each row links to the community report or benchmark database it came from.
Model Quantization Relative Tokens/sec VRAM used Source
gemma-3-4b-it FP16 vllm 3385.7 tok/s 8.1 GB DatabaseMart 2025-03-01
llama2:7b FP16 vllm 2246.0 tok/s vLLM GitHub Discussions 2023-06-27
llama-7b vllm 2246.0 tok/s vllm GitHub Discussion #275 2023-06-27
DeepSeek-R1-Distill-Llama-8B FP16 vllm 2225.3 tok/s 15.0 GB DatabaseMart 2025-03-01
deepseek-r1-distill-llama:8b BF16 vllm 2225.0 tok/s 15.0 GB Databasemart.com 2026-08-25
qwen2.5:7b FP16 vllm 2091.2 tok/s 15.0 GB DatabaseMart 2025-03-01
qwen2.5:7b BF16 vllm 2091.0 tok/s 15.0 GB Databasemart.com 2026-08-25
qwen3:32b FP16 vllm 913.7 tok/s DatabaseMart 2025-05-01
open_llama:13b FP16 vllm 745.2 tok/s vLLM GitHub Discussions 2023-07-03
DeepSeek-R1-Distill-Qwen-14B FP16 vllm 615.6 tok/s 28.0 GB DatabaseMart 2025-03-01
deepseek-r1:14b FP16 vllm 615.6 tok/s 28.0 GB DatabaseMart 2026-08-03
deepseek-r1-distill-qwen:14b BF16 vllm 615.6 tok/s 28.0 GB Databasemart.com 2026-08-25

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA A100 PCIe 40GB — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
PassMark G3D Mark 13,326 points PassMark VideoCardBenchmark 2025-01-01
HPL (High Performance LINPACK) FP64 10,940 Gflops Puget Systems Labs 2021-05-21
Blender (Bizon composite) 3,788 points Bizon Tech GPU Benchmarks 2024-01-01
V-Ray GPU 1,555 points Bizon Tech GPU Benchmarks 2024-01-01
OctaneBench 498 points Bizon Tech GPU Benchmarks 2024-01-01

Full Specifications

tdp w250
vram gb40
cuda cores6912
memory typeHBM2

NVIDIA A100 PCIe 40GB — Frequently Asked Questions

What is the NVIDIA A100 PCIe 40GB best used for?
NVIDIA A100 PCIe 40GB is positioned as a 40 GB VRAM Ampere-family graphics card. Use it for high-end 4K gaming and local LLM inference. See the synthetic + AI benchmark tables below for measured performance.
When was the NVIDIA A100 PCIe 40GB released, and what was its launch MSRP?
NVIDIA A100 PCIe 40GB launched in 2020 at a $12,000 MSRP. Street prices diverge from launch pricing over a product's lifetime — check the linked Amazon listings on this page for current availability.
Where do the benchmark numbers on this page come from?
Synthetic benchmarks are scraped from public databases (TechPowerUp, PassMark, Geekbench Browser, Cinebench leaderboards). AI inference numbers come from the LocalLLaMA community (Reddit threads, llama.cpp / Ollama discussion logs, and Phoronix when available). Every benchmark row carries an inline source citation — click through to verify the original number.
Can the NVIDIA A100 PCIe 40GB run local LLMs?
Yes — NVIDIA A100 PCIe 40GB has 12 AI inference benchmarks on file (see the AI Inference Performance section above for model + tokens-per-second numbers). With 40 GB VRAM, it fits the popular 32B-parameter open-weight models at Q4 quantization comfortably.
Where can I buy the NVIDIA A100 PCIe 40GB?
Active Amazon listings aren't on file for this exact SKU yet. See the linked benchmark sources and the Compare tool for adjacent parts that may be in stock — and check the /benchmarks index for the latest curated picks in this category.

Buying guides that rank the NVIDIA A100 PCIe 40GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA A100 PCIe 40GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →