Skip to main content
NVIDIA Tesla P40 24GB
NVIDIA · GPU · Pascal Pro

NVIDIA Tesla P40 24GB — Benchmarks & Specs

24 GB VRAM250W TDP$6,800 MSRP2016

Bottom line: how fast is the NVIDIA Tesla P40 24GB?

For local LLM inference it generates 29.2 tokens/sec running qwen3-30b-a3b (a mixture-of-experts model with ~3B active parameters, so dense models of similar size run far slower) at q4_K_M under llama.cpp, per LocalScore.ai. In Geekbench OpenCL it scores 62,287 points, per Geekbench Browser / askgeek.io. Its 24 GB of VRAM is the binding constraint for local inference: the qwen3-30b-a3b run above (30.5B parameters) is the largest model on file at 4-bit quantization on this card.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA Tesla P40 24GB is a graphics card from the Pascal Pro family released in 2016 from NVIDIA. Key on-paper specs include 24 GB of GDDR5 VRAM, 250W TDP. It launched with a $6,800 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 10 synthetic benchmark results, 21 community AI inference reports (top 12 shown), compiled from LocalScore.ai, TinyComputers.io, TopCPU, 3DMark, Geekbench Browser, Geekbench Browser / askgeek.io, KnightLi llama.cpp GPU Benchmark Scoreboard, Like2Byte Tesla P40 LLM Guide, llama.cpp GitHub, PassMark Software; each table row links to its source. Read this page when shopping the NVIDIA Tesla P40 24GB, comparing it against other graphics cards in your build, or sizing it for a specific workload (synthetic benchmark scores or local LLM inference).

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA Tesla P40 24GB, in generated tokens per second for a single request (one user's generation speed). Higher is better; each row links to the community report or benchmark database it came from.
Model Quantization Relative Tokens/sec VRAM used Source
llama3.2:3b — ollama 94.3 tok/s — TinyComputers.io 2026-03-01
llama3.2:1b q4_K_M llama.cpp 90.5 tok/s — LocalScore.ai 2025-01-01
llama2:7b q4_0 llama.cpp 54.7 tok/s — llama.cpp GitHub Discussions 2024-12-01
llama-2:7b Q4_0 llama.cpp 54.7 tok/s — llama.cpp GitHub 2025-08-01
llama2:7b q4_0 llama.cpp 54.7 tok/s — KnightLi llama.cpp GPU Benchmark Scoreboard 2026-04-23
qwen2.5:7b — ollama 52.7 tok/s — TinyComputers.io 2026-03-01
llama3.1:8b — ollama 47.8 tok/s — TinyComputers.io 2026-03-01
mistral:7b q4_K_M llama.cpp 45.0 tok/s — Like2Byte Tesla P40 LLM Guide 2026-02-20
llama2:7b q4_0 llama.cpp 40.9 tok/s — LocalScore.ai 2024-06-01
llama-2:7b Q4_0 llama.cpp 40.9 tok/s — LocalScore.ai 2025-01-01
qwen3:30b-a3b q4_K_M llama.cpp 29.2 tok/s — LocalScore.ai 2025-01-01
qwen3-30b-a3b q4_K_M llama.cpp 29.2 tok/s — LocalScore.ai 2025-06-01

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA Tesla P40 24GB — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
Geekbench OpenCL 62,287 points Geekbench Browser / askgeek.io 2024-01-01
Geekbench 5 OpenCL 62,287 points Geekbench Browser 2021-07-16
PassMark G3D Mark 11,596 points PassMark Software 2026-06-09
3DMark Time Spy 8,079 points 3DMark 2023-01-01
3DMark Time Spy 8,079 points TopCPU 2024-01-01
3DMark Time Spy 7,418 points 3DMark (UL Benchmarks) 2020-01-01
PassMark GPU Compute 4,082 Ops/Sec PassMark PerformanceTest 2026-04-30
3DMark Time Spy Extreme 3,894 points TopCPU 2024-01-01
Blender 797 points Blender Benchmark / topcpu.net 2023-01-01
OctaneBench 166 points OctaneBench / topcpu.net 2023-01-01

Full Specifications

tdp w250
vram gb24
cuda cores3840
memory typeGDDR5

NVIDIA Tesla P40 24GB — Frequently Asked Questions

What is the NVIDIA Tesla P40 24GB best used for?
NVIDIA Tesla P40 24GB is a 24 GB VRAM Pascal Pro-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA Tesla P40 24GB released, and what was its launch MSRP?
NVIDIA Tesla P40 24GB launched in 2016 at a $6,800 MSRP. Street prices drift from launch pricing over a product's life, so check the current listing before buying.
Where do the NVIDIA Tesla P40 24GB benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA Tesla P40 24GB run local LLMs?
Yes — AI inference results for the NVIDIA Tesla P40 24GB are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 24 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA Tesla P40 24GB?
SpecPicks has no verified listing of the NVIDIA Tesla P40 24GB on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA Tesla P40 24GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA Tesla P40 24GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →