Skip to main content
NVIDIA L40S 48GB
NVIDIA · GPU · Ada Lovelace Pro

NVIDIA L40S 48GB — Benchmarks & Specs

48 GB VRAM350W TDP$22,000 MSRP2023

Bottom line: how fast is the NVIDIA L40S 48GB?

For local LLM inference it generates 15.3 tokens/sec running llama3:70b at q4_K_M under llama.cpp, per GPU-Benchmarks-on-LLM-Inference. In Geekbench 6 OpenCL it scores 352,507 points, per Tom's Hardware. Its 48 GB of VRAM is the binding constraint for local inference: the llama3:70b run above (70B parameters) is the largest model on file at 4-bit quantization on this card.

Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.

The NVIDIA L40S 48GB is a graphics card from the Ada Lovelace Pro family released in 2023 from NVIDIA. Key on-paper specs include 48 GB of GDDR6 VRAM, 350W TDP. It launched with a $22,000 MSRP, though street prices typically diverge meaningfully from launch pricing. Data on this page draws on 7 synthetic benchmark results, 28 community AI inference reports (top 12 shown), compiled from GPU-Benchmarks-on-LLM-Inference, VMware Cloud Foundation Blog, FPSBench, PassMark Software, Red Hat, technical.city, Crusoe, Fluence, Koyeb, NVIDIA NIM LLMs Benchmarking, Tom's Hardware; each table row links to its source. Read this page when shopping the NVIDIA L40S 48GB, comparing it against other graphics cards in your build, or sizing it for a specific workload (synthetic benchmark scores or local LLM inference).

AI Inference Performance

Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.

Local LLM inference throughput on the NVIDIA L40S 48GB, in generated tokens per second for a single request (one user's generation speed). Higher is better; each row links to the community report or benchmark database it came from. Batched serving runs are listed separately below.
Model Quantization Relative Tokens/sec VRAM used Source
Llama 3 8B Instruct W4A8KV4 (QServe) QServe 3568.8 tok/s — Crusoe 2024-08-02
llama3:8b q4_K_M llama.cpp 113.6 tok/s — GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01
Llama 3 8B Q4_K_M llama.cpp 113.6 tok/s — GPU-Benchmarks-on-LLM-Inference (GitHub) 2024-05-01
llama3.1:8b FP8 vllm 72.6 tok/s — NVIDIA NIM LLMs Benchmarking 2025-03-01

Batched serving throughput

Aggregate output tokens per second on the NVIDIA L40S 48GB summed across many concurrent requests (serving runtimes such as vLLM). This is total server throughput, not one user's generation speed — each request generates far slower — so these rows are not comparable with the single-stream table above.
Model Quantization Relative Tokens/sec (aggregate) VRAM used Source
llama3:8b W4A8KV4 QServe 128 concurrent requests 3568.8 tok/s aggregate — Crusoe AI / MIT Han Lab 2024-07-01
Llama 3.1 8B Instruct FP8 vLLM (MLPerf v5.1 Offline) batched 1642.0 tok/s aggregate — Red Hat (MLPerf Inference v5.1) 2025-09-10
Llama 3.1 8B Instruct FP8 vLLM (MLPerf v5.1 Server) batched 1207.0 tok/s aggregate — Red Hat (MLPerf Inference v5.1) 2025-09-10
llama3.1:8b FP16 vllm 8 concurrent requests 336.0 tok/s aggregate — Koyeb 2024-10-01
Llama 3.1 8B FP16 TensorRT-LLM (batch 8) batched 325.1 tok/s aggregate — Fluence 2025-12-30
mistral:7b-v0.3 BF16 vllm 10 concurrent requests 237.8 tok/s aggregate — VMware Cloud Foundation Blog 2024-09-25
llama3.1:8b BF16 vllm 10 concurrent requests 208.1 tok/s aggregate — VMware Cloud Foundation Blog 2024-09-25
qwen2.5:14b BF16 vllm 10 concurrent requests 113.2 tok/s aggregate — VMware Cloud Foundation Blog 2024-09-25

Synthetic Benchmarks

Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.

Synthetic benchmark scores for the NVIDIA L40S 48GB — higher is better. Each row links to the public database the number was taken from.
Benchmark Relative Score Source
Geekbench 6 OpenCL 352,507 points Tom's Hardware 2024-06-25
Geekbench 5 OpenCL 336,576 points technical.city 2024-06-01
Geekbench OpenCL 330,727 points FPSBench (aggregated from Geekbench Browser) 2024-03-01
Geekbench Vulkan 260,799 points FPSBench (aggregated from Geekbench Browser) 2024-03-01
Geekbench 5 Vulkan 260,799 points technical.city 2024-06-01
PassMark G3D Mark 20,023 points PassMark Software 2024-01-01
PassMark GPU Compute 18,866 Ops/Sec PassMark Software 2025-03-26

Full Specifications

tdp w350
vram gb48
cuda cores18176
memory typeGDDR6

NVIDIA L40S 48GB — Frequently Asked Questions

What is the NVIDIA L40S 48GB best used for?
NVIDIA L40S 48GB is a 48 GB VRAM Ada Lovelace Pro-family graphics card. High-end 4K gaming and local LLM inference are the headline use cases; the benchmark tables on this page place it within its generation.
When was the NVIDIA L40S 48GB released, and what was its launch MSRP?
NVIDIA L40S 48GB launched in 2023 at a $22,000 MSRP. Street prices drift from launch pricing over a product's life, so check the current listing before buying.
Where do the NVIDIA L40S 48GB benchmark numbers come from?
Each benchmark row on this page names its source (public benchmark databases such as TechPowerUp, PassMark, Geekbench and Cinebench, and community reports such as r/LocalLLaMA threads) and links to it, so every number can be checked against the original.
Can the NVIDIA L40S 48GB run local LLMs?
Yes — AI inference results for the NVIDIA L40S 48GB are on file; the AI Inference Performance section on this page lists model, quantization and tokens-per-second for each run. With 48 GB VRAM, 32B-parameter open-weight models fit at Q4 quantization.
Where can I buy the NVIDIA L40S 48GB?
SpecPicks has no verified listing of the NVIDIA L40S 48GB on file yet; the Amazon search link at the top of this page shows its current offers. SpecPicks earns a small affiliate commission on qualifying purchases.

Buying guides that rank the NVIDIA L40S 48GB's class

This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.

Editorial guides covering the NVIDIA L40S 48GB

In-depth SpecPicks reviews, build guides, and head-to-heads referencing this graphics card.

More guides & deep dives from the SpecPicks archive

Browse all articles & guides →

More reviews from the SpecPicks archive

Browse all reviews →

More buying guides from SpecPicks

Browse all buying guides →

Hardware benchmark data on SpecPicks

All benchmarks →