NVIDIA L40 48GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA L40 48GB?
For local LLM inference it generates 15.3 tokens/sec running llama3.1:70b at Q4_K_M under llama.cpp, per GPU-Benchmarks-on-LLM-Inference. In Geekbench 5 OpenCL it scores 333,904 points, per technical.city. Its 48 GB of VRAM is the binding constraint for local inference: the llama3.1:70b run above (70B parameters) is the largest model on file at 4-bit quantization on this card.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA L40 48GB is a graphics card from the Ada Lovelace Pro family released in 2023 from NVIDIA. Key on-paper specs include 48 GB of GDDR6 VRAM, 300W TDP. It launched with a $18,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 7 synthetic benchmark results, 11 community AI inference reports, compiled from CpuTronic, GPU-Benchmarks-on-LLM-Inference, LocalScore, technical.city, Crusoe AI Blog, Koyeb GPU Benchmarks, llama.cpp GitHub Discussion #15013, PassMark Software, vLLM GitHub Issues #5007, VMware Cloud Foundation Blog; each table row links to its source. Read this page when shopping the NVIDIA L40 48GB, comparing it against other graphics cards in your build, or sizing it for a specific workload (synthetic benchmark scores or local LLM inference).
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| llama2:7b | q4_0 llama.cpp | 152.0 tok/s | — | llama.cpp GitHub Discussion #15013 2024-01-01 | |
| llama3:8b | q4_K_M llama.cpp | 113.6 tok/s | — | XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01 | |
| llama3:8b | FP16 vllm | 54.1 tok/s | — | VMware Cloud Foundation Blog 2024-09-25 | |
| llama-3.1:8b | — vllm | 46.0 tok/s | — | Koyeb GPU Benchmarks 2024-06-01 | |
| llama3.1:8b | q4_K_M llamafile | 45.1 tok/s | — | LocalScore 2025-04-01 | |
| llama3:8b | FP16 llama.cpp | 43.4 tok/s | — | XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01 | |
| qwen2.5:14b | q4_K_M llamafile | 24.7 tok/s | — | LocalScore 2025-04-01 | |
| llama3:70b | q4_K_M llama.cpp | 15.3 tok/s | — | XiongjieDai/GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01 | |
| llama3.1:70b | Q4_K_M llama.cpp | 15.3 tok/s | — | GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-07-01 |
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| llama3:8b | W4A8KV4 QServe 128 concurrent requests | 3568.8 tok/s aggregate | — | Crusoe AI Blog 2024-08-02 | |
| llama-2:7b | FP16 vllm batched | 3115.3 tok/s aggregate | — | vLLM GitHub Issues #5007 2024-05-23 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| Geekbench 5 OpenCL | 333,904 points | technical.city (Geekbench Browser community) 2024-01-01 | |
| Geekbench 5 OpenCL | 292,357 points | CpuTronic 2023-10-01 | |
| Geekbench 5 Vulkan | 249,130 points | CpuTronic 2023-10-01 | |
| Geekbench 5 Vulkan | 237,295 points | technical.city (Geekbench Browser community) 2024-01-01 | |
| PassMark G3D Mark | 27,355 points | PassMark Software 2025-12-11 | |
| PassMark GPU Compute | 23,067 points | PassMark VideoCardBenchmark 2023-07-01 | |
| Blender Benchmark | 4,336 points | CpuTronic 2023-10-01 |
Full Specifications
| tdp w | 300 |
|---|---|
| vram gb | 48 |
| cuda cores | 18176 |
| memory type | GDDR6 |
NVIDIA L40 48GB — Frequently Asked Questions
What is the NVIDIA L40 48GB best used for?
When was the NVIDIA L40 48GB released, and what was its launch MSRP?
Where do the NVIDIA L40 48GB benchmark numbers come from?
Can the NVIDIA L40 48GB run local LLMs?
Where can I buy the NVIDIA L40 48GB?
Buying guides that rank the NVIDIA L40 48GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA L40 48GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Ryzen 5 5600X vs Ryzen 5 5600G for CPU-Only Gemma 3 12B
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- Best Hardware for Local OCR and Document AI in 2026
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Prime Big Deal Days 2026 GPU Deals for Local LLMs: 12GB vs 16GB Cards
- Best GPU for gpt-oss 20B in 2026
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- Best 1440p Gaming GPUs in 2026
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- How to Build a Windows 98 Retro PC in 2026
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- Best Retro Handhelds in 2026 — From $35 to $500
More reviews from the SpecPicks archive
Browse all reviews →- Best GPU for the Ryzen 7 5800X3D in 2026: No-Bottleneck Picks
- Raspberry Pi OS Moves to Linux 6.18 LTS — Real Performance Gains
- 8BitDo Pro 2 vs GameSir G7 SE vs DualSense: the best PC controller for retro emulation in 2026
- Voodoo 5 6000 in 2026: The Card That Never Shipped, Reproduced
- GeForce FX 5900 Won't POST or Crashes in WinXP: Troubleshooting Guide
- How to run Qwen 3 32B on Apple M4 Pro
- AMD Ryzen AI Halo vs NVIDIA DGX Spark: Local-AI Mini-Box Showdown
- Building a New Gaming PC in 2026: The Complete Guide
- Forza Horizon 6 Advanced Shader Delivery: 4-Second Loads vs 90 Seconds Explained
- Best SATA SSD for Boot Drive Upgrades on Aging Desktops (2026)
- Ideogram 4.0 Open Weights on an RTX 3060 12GB: Local Text-to-Image in 2026
- Pro-Level 180Hz 1440p Gaming Monitor for $159: Gigabyte Deal
- Two GTX 1050 Ti 4GB vs One RTX 3060 12GB for 8B Local LLMs
- vLLM 0.21 Adds Intel GPU Support: What It Means for Budget AI Rigs
- Best Budget Local-AI Workstation Parts in 2026: 5 Picks
- CompactFlash + IDE Storage for a Period-Correct Windows 98 Build
- Best SSD for a Retro PC Build: CompactFlash vs SATA in 2026
- Local 13B LLM Inference on a $700 Used Build: Ryzen 7 3700X + RTX 3060 12GB Benchmarked
- RE Requiem and the AAA Games That Prove Steam Deck
- RTX 3060 12GB: Ollama vs llama.cpp vs vLLM Token Speed (2026)
- Dual RTX 3090 PC Builds: PSU, Cooling, and Cost in 2025
- Using LLMs to Install Vintage GPU Drivers on Win98 and WinXP: A Field Report from Our Retro-Agent Fleet
- Best Raspberry Pi for a Retro Console Build in 2026: Pi 4 8GB vs Pi Zero 2 W
- Best Plug-and-Play Retro Gaming Consoles to Buy in 2026
More buying guides from SpecPicks
Browse all buying guides →- Best CPU Coolers for 2026
- Best GPUs for Running Local LLMs in 2026
- Best Gaming Monitors for 2026
- Best NVMe External Enclosures for 2026
- Best NVMe SSDs for Gaming in 2026
- Best External SSDs for Content Creators in 2026
- Best Gaming Mice for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best CPUs for Content Creators in 2026
- Best PC Cases for Building in 2026
- Best CPUs for Gaming in 2026
- Best AM5 Motherboards for 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best Controllers for PC Gaming in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best GPUs for 4K Gaming in 2026
- Best 4K Monitors for Content Creators in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best Graphics Cards for Gaming in 2026
- Best Mechanical Keyboards for Gaming in 2026
Hardware benchmark data on SpecPicks
All benchmarks →- Apple M4 8 Core — benchmarks & specs
- DDR5-6400 CL32 32GB (2x16) — benchmarks & specs
- AMD Ryzen 9 5980HS — benchmarks & specs
- Radeon RX 7900M — benchmarks & specs
- Ryzen 5 4500U — benchmarks & specs
- Quadro RTX 5000 (Mobile) — benchmarks & specs
- AMD Ryzen 5 7235HS — benchmarks & specs
- AMD Ryzen 9 5900XT — benchmarks & specs
- Apple M3 Pro 11 Core — benchmarks & specs
- Intel Core i5-14500 — benchmarks & specs
- Ryzen 5 3500U — benchmarks & specs
- Intel Arc B390 GPU — benchmarks & specs
- Pentium 4 2.53GHz (Northwood) — benchmarks & specs
- AMD Ryzen Threadripper 9960X 24-Cores — benchmarks & specs
- GeForce RTX 4090 D — benchmarks & specs
- AMD Ryzen 7 7840S — benchmarks & specs
- Apple M1 Ultra 20 Core — benchmarks & specs
- NVIDIA B200 — benchmarks & specs
- AMD Ryzen 5 5600F — benchmarks & specs
- AMD Ryzen 9 7945HX3D — benchmarks & specs
- NVIDIA RTX A2000 12GB — benchmarks & specs
- AMD Ryzen Threadripper PRO 7965WX — benchmarks & specs
- Radeon RX 6550S — benchmarks & specs
- Radeon RX 9070 GRE — benchmarks & specs