NVIDIA L40S 48GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA L40S 48GB?
For local LLM inference it generates 15.3 tokens/sec running llama3:70b at q4_K_M under llama.cpp, per GPU-Benchmarks-on-LLM-Inference. In Geekbench 6 OpenCL it scores 352,507 points, per Tom's Hardware. Its 48 GB of VRAM is the binding constraint for local inference: the llama3:70b run above (70B parameters) is the largest model on file at 4-bit quantization on this card.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA L40S 48GB is a graphics card from the Ada Lovelace Pro family released in 2023 from NVIDIA. Key on-paper specs include 48 GB of GDDR6 VRAM, 350W TDP. It launched with a $22,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 7 synthetic benchmark results, 28 community AI inference reports (top 12 shown), compiled from GPU-Benchmarks-on-LLM-Inference, VMware Cloud Foundation Blog, FPSBench, PassMark Software, Red Hat, technical.city, Crusoe, Fluence, Koyeb, NVIDIA NIM LLMs Benchmarking, Tom's Hardware; each table row links to its source. Read this page when shopping the NVIDIA L40S 48GB, comparing it against other graphics cards in your build, or sizing it for a specific workload (synthetic benchmark scores or local LLM inference).
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| Llama 3 8B Instruct | W4A8KV4 (QServe) QServe | 3568.8 tok/s | — | Crusoe 2024-08-02 | |
| llama3:8b | q4_K_M llama.cpp | 113.6 tok/s | — | GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01 | |
| Llama 3 8B | Q4_K_M llama.cpp | 113.6 tok/s | — | GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01 | |
| llama3.1:8b | FP8 vllm | 72.6 tok/s | — | NVIDIA NIM LLMs Benchmarking 2025-03-01 |
Batched serving throughput
| Model | Quantization | Relative | Tokens/sec (aggregate) | VRAM used | Source |
|---|---|---|---|---|---|
| llama3:8b | W4A8KV4 QServe 128 concurrent requests | 3568.8 tok/s aggregate | — | Crusoe AI / MIT Han Lab 2024-07-01 | |
| Llama 3.1 8B Instruct | FP8 vLLM (MLPerf v5.1 Offline) batched | 1642.0 tok/s aggregate | — | Red Hat (MLPerf Inference v5.1) 2025-09-10 | |
| Llama 3.1 8B Instruct | FP8 vLLM (MLPerf v5.1 Server) batched | 1207.0 tok/s aggregate | — | Red Hat (MLPerf Inference v5.1) 2025-09-10 | |
| llama3.1:8b | FP16 vllm 8 concurrent requests | 336.0 tok/s aggregate | — | Koyeb 2024-10-01 | |
| Llama 3.1 8B | FP16 TensorRT-LLM (batch 8) batched | 325.1 tok/s aggregate | — | Fluence 2025-12-30 | |
| mistral:7b-v0.3 | BF16 vllm 10 concurrent requests | 237.8 tok/s aggregate | — | VMware Cloud Foundation Blog 2024-09-25 | |
| llama3.1:8b | BF16 vllm 10 concurrent requests | 208.1 tok/s aggregate | — | VMware Cloud Foundation Blog 2024-09-25 | |
| qwen2.5:14b | BF16 vllm 10 concurrent requests | 113.2 tok/s aggregate | — | VMware Cloud Foundation Blog 2024-09-25 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| Geekbench 6 OpenCL | 352,507 points | Tom's Hardware 2024-06-25 | |
| Geekbench 5 OpenCL | 336,576 points | technical.city 2024-06-01 | |
| Geekbench OpenCL | 330,727 points | FPSBench (aggregated from Geekbench Browser) 2024-03-01 | |
| Geekbench Vulkan | 260,799 points | FPSBench (aggregated from Geekbench Browser) 2024-03-01 | |
| Geekbench 5 Vulkan | 260,799 points | technical.city 2024-06-01 | |
| PassMark G3D Mark | 20,023 points | PassMark Software 2024-01-01 | |
| PassMark GPU Compute | 18,866 Ops/Sec | PassMark Software 2025-03-26 |
Full Specifications
| tdp w | 350 |
|---|---|
| vram gb | 48 |
| cuda cores | 18176 |
| memory type | GDDR6 |
NVIDIA L40S 48GB — Frequently Asked Questions
What is the NVIDIA L40S 48GB best used for?
When was the NVIDIA L40S 48GB released, and what was its launch MSRP?
Where do the NVIDIA L40S 48GB benchmark numbers come from?
Can the NVIDIA L40S 48GB run local LLMs?
Where can I buy the NVIDIA L40S 48GB?
Buying guides that rank the NVIDIA L40S 48GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA L40S 48GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.
- Running Llama 3.1 70B Locally: Hardware Requirements & Performance Benchmarks
- NVIDIA RTX A6000 48GB Review: The Workstation Card That Still Owns Local 70B Inference (2026)
- NVIDIA RTX PRO 6000 Blackwell vs RTX A6000: Is 96 GB Worth $4,000 More?
- Prime Big Deal Days 2026: Best Mini PC Deals for Local AI and Home Labs
- Raspberry Pi 4 8GB vs Ryzen 5 2600 for Gemma 3 4B: Which Cheap Box Wins?
- Prime Big Deal Days 2026 GPU Deals for Local LLMs: 12GB vs 16GB Cards
- Ryzen 5 2600 vs Ryzen 5 5600X for CPU-Only gpt-oss 20B (2026)
- Arc Pro B60 Dual 48GB vs Dual RTX 3060 12GB for 70B Local LLMs (2026)
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
- Best Retro Handhelds in 2026 — From $35 to $500
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- How to Build a Windows 98 Retro PC in 2026
- Best 1440p Gaming GPUs in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Claude Opus 4.8 Tops the Intelligence Index: Cloud vs Local on a 3060
- Acti Brings AI Agents Into Your Smartphone Keyboard
- Cloning a Win98 Voodoo3 Boot Drive: SATA-IDE Adapter Workflow
- Best Mid-Range CPU for Streaming and Gaming in 2026: Ryzen 7 5800X vs 5700X vs 5600G
- GameSir G7 SE vs 8BitDo Pro 2 vs DualSense for PC
- Microsoft's SkillOpt Boosts GPT-5.5 With Just a Trained Markdown File
- Best GPU for Forza Horizon 6 at 1440p in 2026 (Without Breaking the Bank)
- Intel Arc Pro B60 AI: 24GB VRAM for Local LLMs in 2026
- Is the KOORUI 27-Inch 4K QD-Mini LED a Good Match for an RTX 3060?
- Sound Blaster Audigy 2 ZS Won't Install on Windows XP: Driver Hunt + Fix
- Best Gaming Monitor for Console + PC in 2026
- Voodoo3 3500 TV on Windows 98 SE: AI-Assisted Driver Install Walkthrough in 2026
- Raspberry Pi AI HAT+ (26 TOPS): What It Actually Runs in 2026
- Best Controller for PC and Steam Gaming in 2026: G7 SE vs 8BitDo Pro 2 vs DualSense
- OpenAI Codex Now Records and Replays Your Workflow: the Local-Rig Angle
- Best PC Streaming Starter Kit in 2026: 5 Picks
- Build a Fully Local PDF-to-Audiobook Pipeline on Jetson Orin Nano Super
- Best Budget Streaming Mic + Capture Setup: Blue Yeti vs HyperX QuadCast 2 S
- Sound Blaster in 2026: Using the Creative Sound BlasterX G6 on a Retro and Modern PC
- Best 1440p Gaming Monitor in 2026: 5 Picks for Esports, RPGs, and PS5
- Best Budget Streaming & Podcast Gear in 2026
- Best Streaming Gear for Twitch Beginners in 2026
- Best 4K Monitor for Forza Horizon 6 on PC (2026)
- Solar-Powered Bird Identifier: Raspberry Pi Zero 2W + AI Camera
More buying guides from SpecPicks
Browse all buying guides →- Best 1440p 240Hz Gaming Monitors in 2026
- Best CPUs for Content Creators in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best Gaming Mice for 2026
- Best PC Cases for Building in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best Gaming Monitors for 2026
- Best 4K Monitors for Content Creators in 2026
- Best NVMe SSDs for Gaming in 2026
- Best Controllers for PC Gaming in 2026
- Best CPU Coolers for 2026
- Best CPUs for Gaming in 2026
- Best GPUs for Running Local LLMs in 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best External SSDs for Content Creators in 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best GPUs for 4K Gaming in 2026
- Best Graphics Cards for Gaming in 2026
- Best AM5 Motherboards for 2026
- Best NVMe External Enclosures for 2026
- Best Mechanical Keyboards for Gaming in 2026
Hardware benchmark data on SpecPicks
All benchmarks →- Apple M3 Pro 11 Core — benchmarks & specs
- AMD Ryzen 9 7845HX — benchmarks & specs
- Radeon RX 7700S — benchmarks & specs
- RTX 5000 Ada Generation — benchmarks & specs
- Ryzen 7 7800X3D — benchmarks & specs
- NVIDIA GeForce RTX 3060 — benchmarks & specs
- NVIDIA GeForce RTX 5080 — benchmarks & specs
- Ryzen 7 7735HS — benchmarks & specs
- Ryzen 5 5500X3D — benchmarks & specs
- Radeon RX 6750 XT — benchmarks & specs
- AMD Ryzen 5 PRO 5650U — benchmarks & specs
- NVIDIA GeForce RTX 5090 — benchmarks & specs
- Ryzen 7 9700X — benchmarks & specs
- AMD Ryzen Threadripper PRO 9985WX — benchmarks & specs
- AMD Ryzen 5 PRO 5655GE — benchmarks & specs
- NVIDIA Jetson Orin Nano 8GB — benchmarks & specs
- AMD Ryzen 5 PRO 5650G — benchmarks & specs
- Ryzen 5 3600 — benchmarks & specs
- Ryzen 9 5950X — benchmarks & specs
- Samsung 990 PRO 2TB — benchmarks & specs
- NVIDIA GeForce GTX 1660 SUPER — benchmarks & specs
- GeForce RTX 3050 6 GB — benchmarks & specs
- NVIDIA GeForce RTX 2080 — benchmarks & specs
- GeForce RTX 6090 — benchmarks & specs