NVIDIA Jetson AGX Orin 64GB — Benchmarks & Specs
Bottom line: how fast is the NVIDIA Jetson AGX Orin 64GB?
For local LLM inference it generates 300.0 tokens/sec running Mistral-7B at unspecified under llama.cpp (CUDA 12.9), per NVIDIA Developer Forums. In MLPerf v3.1 ResNet-50 Image Classification Offline it scores 6,423.63 samples/s, per NVIDIA Developer.
Every figure above is a row in the tables below, and each row links out to the review or public benchmark database the number was taken from. SpecPicks aggregates published measurements; it does not report first-party benchmark runs.
The NVIDIA Jetson AGX Orin 64GB is an APU from the Jetson Orin family released in 2022 from NVIDIA. It launched with a $1,999 MSRP, though street prices typically diverge meaningfully from launch pricing — see the linked product cards below for current Amazon listings. Data on this page draws on 7 synthetic benchmark results, 12 community AI inference reports, aggregated from public benchmark databases (TechPowerUp, PassMark, Geekbench, Cinebench) and the LocalLLaMA community. Read this page when shopping the NVIDIA Jetson AGX Orin 64GB, comparing it against other APUs in your build, or sizing it for a specific workload (gaming at 1080p/1440p/4K, productivity benchmarks, or local LLM inference).
AI Inference Performance
Tokens per second under each model + quantization. Higher = faster generation. Bars compare runs across the same model.
| Model | Quantization | Relative | Tokens/sec | VRAM used | Source |
|---|---|---|---|---|---|
| Mistral-7B | unspecified llama.cpp (CUDA 12.9) | 300.0 tok/s | — | NVIDIA Developer Forums 2026-01-24 | |
| qwen2.5-vl:3b | — tensorrt-llm | 225.7 tok/s | — | NVIDIA Developer Forums 2025-03-01 | |
| deepseek-r1-distill-qwen:7b | — vllm | 180.4 tok/s | — | NVIDIA Developer Forums 2025-02-01 | |
| gemma2:2b | FP16 tensorrt-llm | 120.0 tok/s | — | IoT Digital Twin PLM 2026-06-01 | |
| gemma2:2b | FP16 TensorRT-LLM | 120.0 tok/s | — | IoT Digital Twin PLM 2026-04-24 | |
| llama3.2:3b | INT4 TensorRT-LLM | 85.0 tok/s | — | IoT Digital Twin PLM 2026-04-24 | |
| llama3.2:3b | INT4 tensorrt-llm | 85.0 tok/s | 3.2 GB | IoT Digital Twin PLM 2026-06-01 | |
| phi3.5:mini | INT4 tensorrt-llm | 62.0 tok/s | — | IoT Digital Twin PLM 2026-06-01 | |
| phi3.5:mini | INT4 TensorRT-LLM | 62.0 tok/s | — | IoT Digital Twin PLM 2026-04-24 | |
| llama3.1:8b | q4_K_M llama.cpp | 52.0 tok/s | — | ProventusNova Blog 2026-06-10 | |
| phi3:mini | q4_K_M llama.cpp | 47.0 tok/s | 2.8 GB | MultimodalFlow Blog 2026-05-01 | |
| llama2:7b | q4f16_ft mlc | 46.9 tok/s | — | jetson-containers (NVIDIA/dusty-nv) 2024-04-01 |
Synthetic Benchmarks
Higher is better. Bars are scaled within each benchmark family (multi-thread, single-thread, etc.) so you can compare like-with-like at a glance.
| Benchmark | Relative | Score | Source |
|---|---|---|---|
| MLPerf v3.1 ResNet-50 Image Classification Offline | 6,423.63 samples/s | NVIDIA Developer 2023-11-01 | |
| MLPerf v3.1 RNN-T Speech-to-Text Offline | 1,169.98 samples/s | NVIDIA Developer 2023-11-01 | |
| MLPerf v3.1 BERT-Large NLP Offline | 553.69 samples/s | NVIDIA Developer 2023-11-01 | |
| MLPerf Inference v4.0 Edge - YOLOv8n | 383 FPS | lowtouch.ai 2025-02-27 | |
| YOLOv8l (TensorRT, 60W) | 95 FPS | lowtouch.ai 2025-02-27 | |
| MLPerf v4.0 GPT-J 6B Summarization Offline | 0.15 samples/s | NVIDIA Developer 2024-10-01 | |
| GPT-J 6B (MLPerf) | 0.15 samples/sec | lowtouch.ai 2025-02-27 |
Full Specifications
| ram gb | 64 |
|---|---|
| ai tops | 275 |
| cpu cores | 12 |
| gpu cores | 2048 |
NVIDIA Jetson AGX Orin 64GB — Frequently Asked Questions
What is the NVIDIA Jetson AGX Orin 64GB best used for?
When was the NVIDIA Jetson AGX Orin 64GB released, and what was its launch MSRP?
Where do the benchmark numbers on this page come from?
How does the NVIDIA Jetson AGX Orin 64GB compare to its predecessor?
Where can I buy the NVIDIA Jetson AGX Orin 64GB?
Buying guides that rank the NVIDIA Jetson AGX Orin 64GB's class
This page is the raw performance data. The guides below turn it into a ranked pick for a specific build.
Editorial guides covering the NVIDIA Jetson AGX Orin 64GB
In-depth SpecPicks reviews, build guides, and head-to-heads referencing this APU.
- Arc B580 12GB vs RTX 3060 12GB: Which 12GB Card for 1440p in 2026?
- Best Wired Gear for Tournament-Legal LAN Play in 2026
- Alien: Isolation on Steam Deck: Best Settings for 2025
- The Greatest GPU Awards: A Decade of Standout Cards
- GTX 1660 VRAM: Why 8GB Doesn't Exist (Real Specs)
- Best Gear for Online D&D and Virtual Tabletop Nights in 2026
- Logitech G29 vs HORI Racing Wheel Overdrive: Which Entry Sim Wheel Wins?
- How to Filter Steam Deck's 20K+ Games by FPS, Price, HLTB
- Browse all SpecPicks reviews →
More guides & deep dives from the SpecPicks archive
Browse all articles & guides →- How to Build a Windows 98 Retro PC in 2026
- The Complete Voodoo5 5500 AGP Driver Guide (2026 Edition)
- Best Budget Gaming PC Build 2026 — ~$1,000 ($800 on Sale)
- Best Retro Handhelds in 2026 — From $35 to $500
- Emulation Hardware in 2026: FPGA, Software, and Cart-Reader Ecosystems
- Best 1440p Gaming GPUs in 2026
- RTX 4070 Super vs RX 7800 XT — Which to Buy in 2026
More reviews from the SpecPicks archive
Browse all reviews →- Open WebUI + Ollama on an RTX 3060: The Self-Hosted ChatGPT Alternative for 2026
- Qwen 3.6 27B with MTP: 2.5x Throughput on Local Hardware (Real Benchmarks)
- Sound Blaster AWE32 vs AWE64: The 1998 MIDI Decision
- Ryzen 5 2600 vs Core i7-9700K for a First 1080p Gaming Build (2026)
- DeepSeek on the US Entity List: What It Means for Local Inference
- StarCraft Brood War in 2026: Joining Battle.net Classic, the Fish Server, and ASL Pro Circuit
- Hobbyist Turns a Raspberry Pi and ADS-B Into a Viral Real-Time Airport Tracker
- CompactFlash as a Win98 Boot Drive: Transcend CF133 + IDE Adapter Walkthrough
- RTX 5090 vs RTX 6000: Specs, Gaming, AI Compared
- vLLM 0.21 Adds Intel GPU Support: What It Means for Budget AI Rigs
- Best CPU for a Budget AI + Gaming Rig: Ryzen 7 5700X vs 5800X vs 5600G
- DeepSeek V4 Pro at $0.04 a Task: When Local Still Beats the Cloud
- Per-Model Hardware Guide: Matching Llama, DeepSeek & Qwen to Your GPU
- Best Raspberry Pi 5 Home Lab Cluster Setup for Self-Hosting (2026)
- Intel Arc Pro B60 Gaming Performance: 2025 Benchmarks
- Razer Blade 16 (2026) Review: Gaming Performance and Endurance
- Is the RTX 3060 12GB Still Worth It for 1440p Gaming in 2026?
- Best Raspberry Pi Alternative in 2025: Full Buying Guide
- Is the RTX 3060 12GB Still the Best Budget GPU for 1080p Esports in 2026?
- NVIDIA GeForce GTX 10 Pascal GPUs: 10th Anniversary
- Etched's Transformer-Only Inference Chip vs Your GPU: What Changes for Local Builders
- vLLM vs llama.cpp on a 12GB GPU: Which Serves Local LLMs Faster?
- AMD Ryzen AI Max 400 'Gorgon Halo': 192GB Unified Memory APU Hits $3,999
- Best GPU for Training CNNs at Home in 2026: The RTX 3060 12GB Case
More buying guides from SpecPicks
Browse all buying guides →- Best NVMe SSDs for Gaming in 2026
- Best AM5 Motherboards for 2026
- Best CPUs for Content Creators in 2026
- Best GPUs for 4K Gaming in 2026
- Best Controllers for PC Gaming in 2026
- Best Tools for Building and Repairing Retro PCs in 2026
- Best CPU Coolers for 2026
- Best PC Cases for Building in 2026
- Best NVMe External Enclosures for 2026
- Best GPU for Running 27B-32B Local LLMs in 2026
- Best 4K Monitors for Content Creators in 2026
- Best Gaming Mice for 2026
- Best Gaming Monitors for 2026
- Best DDR5 RAM for Gaming PCs in 2026
- Best Retro Gaming Consoles & Handhelds for 2026
- Best Graphics Cards for Gaming in 2026
- Best 1440p 240Hz Gaming Monitors in 2026
- Best GPUs for Running Local LLMs in 2026
- Best External SSDs for Content Creators in 2026
- Best CPUs for Gaming in 2026
- Best Mechanical Keyboards for Gaming in 2026