---
格式版本: 2
标题: "Rack-Scale Agentic AI Supercomputer | NVIDIA Vera Rubin NVL72"
原文链接: "https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72"
发布日期: "2026-05-31"
发现时间: "2026-05-22T00:00:00+08:00"
入库时间: "2026-05-28T02:39:38+08:00"
来源平台: "历史资料迁移"
搜索渠道: "legacy_migration"
搜索词: "NVIDIA Vera Rubin rack scale AI、Google A5X NVL72 architecture design manufacturer after:2026-05-18、site:nvidia.com/en-us/data-center NVL72 OR \"AI factory\" OR \"rack\" after:2026-05-18"
匹配关键词:
  - "AI Rack"
  - "Scale-up"
  - "NVIDIA Vera Rubin rack scale AI"
  - "Google A5X NVL72 architecture design manufacturer after:2026-05-18"
  - "site:nvidia.com/en-us/data-center NVL72 OR \"AI factory\" OR \"rack\" after:2026-05-18"
  - "待定数据"
  - "厂家"
  - "Nvidia"
  - "Rack"
  - "Scale"
  - "Agentic"
  - "AI"
  - "Supercomputer"
  - "NVIDIA"
  - "Vera"
  - "Rubin"
  - "NVL72"
  - "22.md"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Firecrawl 原始抓取 + qwen3.6-flash 原文清洗"
清洗工具: "历史资料原文保留 + LLM 正文裁剪"
关联判断: "待人工确认"
关联置信度: 0
关联理由: "历史资料迁移，未重新调用模型判定"
AI优质: "是"
AI打分: 100
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "英伟达官方发布下一代Vera Rubin NVL72整机柜AI超算，明确72卡/36CPU架构、NVLink 6、HBM4、CPO光互连及详细算力/带宽参数，已量产并支持80+生态伙伴，技术细节与商业信号极强，完全契合超节点/AI Rac…"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-06-03T00:28:39+08:00"
AI主题相关性: 20
AI来源权威性: 15
AI新颖性: 20
AI技术细节: 20
AI商业部署信号: 15
AI完整性: 10
采集批次: "2026年5月28日2点39分33秒"
采集批次ID: "legacy-migration-20260528023933"
去重键: "https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72"
---

![Single Rack NVIDIA Vera Rubin NVL72](./assets/01-www.nvidia.com.jpg)
# NVIDIA Vera Rubin NVL72
Building the next frontier of AI.
Overview

## Seven New Chips, One AI Supercomputer
NVIDIA Vera Rubin NVL72 unifies leading-edge technologies from NVIDIA—72 Rubin GPUs, 36 Vera CPUs, ConnectX®-9 SuperNIC™s, and BlueField®-4 DPUs. It scales up intelligence in a rack-scale platform with the NVIDIA NVLink™ 6 switch and scales out with [NVIDIA Quantum-X800 InfiniBand](https://www.nvidia.com/en-us/networking/products/infiniband/quantum-x800/) and Spectrum-X™ Ethernet to power the AI industrial revolution at scale. When deployed with NVIDIA Groq 3 LPX racks, Vera Rubin NVL72 delivers a new class of inference performance for trillion-parameter models and million-token context. Vera Rubin NVL72 is built on the third-generation [NVIDIA MGX™ NVL72 rack](https://www.nvidia.com/en-us/data-center/gb200-nvl72/) design, offering a seamless transition from prior generations. It delivers AI training with one-fourth the GPUs and AI inference at one-tenth the cost per million tokens versus NVIDIA Blackwell. Featuring cable‑free modular tray designs and support from over 80 MGX ecosystem partners, the rack-scale AI supercomputer delivers world‑class performance with rapid deployment.

### NVIDIA Kicks Off the Next Generation of AI With Rubin
The leading-edge platform scales mainstream adoption, slashing cost per token with five breakthroughs for reasoning and agentic AI models.

### NVIDIA Vera Rubin Opens the Agentic AI Frontier
The NVIDIA Vera Rubin platform offers seven new chips, now in full production, to scale the world’s largest AI factories.

Performance

## Massive Efficiency Gains in AI Training and Inference

### Boosting Training Efficiency
NVIDIA Rubin trains mixture-of-expert (MoE) models with one-fourth the number of GPUs over the NVIDIA Blackwell architecture. Projected performance subject to change. Number of GPUs based on a 10T MoE model trained on 100T tokens in a fixed timeframe of 1 month. LLM inference performance subject to change. Cost per 1 million tokens based on Kimi-K2-Thinking model using 32K/8K ISL/OSL comparing Blackwell NVL72 and Rubin NVL72.

### Driving Down Inference Costs
NVIDIA Rubin delivers one-tenth the cost per million tokens compared to NVIDIA Blackwell for highly interactive, deep reasoning agentic AI.

Technology Breakthroughs

## Inside the AI Supercomputer
![NVIDIA Vera Rubin Platform Seven Chips](./assets/02-www.nvidia.com.jpeg)
![NVIDIA Rubin GPU](./assets/03-www.nvidia.com.jpeg)

### NVIDIA Rubin GPU
Rubin GPUs with HBM4 and 50 PF NVFP4 Transformer Engine made for the next generation of AI.

![NVIDIA Vera CPU](./assets/04-www.nvidia.com.jpeg)

### NVIDIA Vera CPU
Vera CPUs are purpose-built for data movement and agentic reasoning, delivering high-bandwidth, energy-efficient compute with deterministic performance.

![NVIDIA NVLink 6 Switch](./assets/05-www.nvidia.com.jpeg)

### NVIDIA NVLink 6 Switch
NVLink 6 switches feature 3.6 terabytes per second (TB/s) of all-to-all, scale-up bandwidth per GPU, enabling high-speed GPU-to-GPU communications for AI.

![NVIDIA ConnectX-9 SuperNIC](./assets/06-www.nvidia.com.jpeg)

### NVIDIA ConnectX-9 SuperNIC
ConnectX‑9 SuperNICs deliver 1.6 terabits per second (Tb/s) of per-GPU bandwidth, with programmable remote direct-memory access (RDMA) for low‑latency, GPU‑direct networking at massive scale.

![NVIDIA BlueField-4 DPU](./assets/07-www.nvidia.com.jpeg)

### NVIDIA BlueField-4 DPU
BlueField-4 DPUs accelerate data processing across storage, networking, cybersecurity, and elastic scaling in AI factories.

![NVIDIA Spectrum-X Ethernet Co-Packaged Optics](./assets/08-www.nvidia.com.jpeg)

### NVIDIA Spectrum-X Ethernet Co-Packaged Optics
Spectrum‑X Ethernet scale‑out switches with integrated silicon photonics deliver 5x better power efficiency, 10x higher network resiliency, and up to 5x more uptime over traditional networking with pluggable transceivers.

![NVIDIA Groq 3 LPU](./assets/09-www.nvidia.com.jpeg)

### NVIDIA Groq 3 LPU
This is the inference accelerator for NVIDIA Vera Rubin NVL72, designed to meet the low-latency and large-context demands of agentic systems. The NVIDIA Groq 3 LPX rack features 256 LPUs with 128 GB SRAM, 40 PB/s memory bandwidth, and 640 TB/s scale-up bandwidth per rack. It is co-designed with Vera Rubin NVL72 to deliver 35x inference performance per watt and up to 10x more revenue opportunity for trillion parameter models relative to Blackwell.

Specifications¹

## NVIDIA Vera Rubin NVL72 Specs
| | NVIDIA Vera Rubin NVL72 | NVIDIA Vera Rubin Superchip | NVIDIA Rubin GPU |
| --- | --- | --- | --- |
| Configuration | 72 NVIDIA Rubin GPUs \| 36 NVIDIA Vera CPUs | 2 NVIDIA Rubin GPUs \| 1 NVIDIA Vera CPU | 1 NVIDIA Rubin GPU |
| NVFP4 Inference | 3,600 PFLOPS | 100 PFLOPS | 50 PFLOPS |
| NVFP4 Training² | 2,520 PFLOPS | 70 PFLOPS | 35 PFLOPS |
| FP8/FP6 Training² | 1,260 PFLOPS | 35 PFLOPS | 17.5 PFLOPS |
| INT8² | 18 POPS | 0.5 POPS | 0.25 POPS |
| FP16/BF16² | 288 PFLOPS | 8 PFLOPS | 4 PFLOPS |
| TF32² | 144 PFLOPS | 4 PFLOPS | 2 PFLOPS |
| FP32 | 9,360 TFLOPS | 260 TFLOPS | 130 TFLOPS |
| FP64 | 2,400 TFLOPS | 67 TFLOPS | 33 TFLOPS |
| FP32 SGEMM³ | 28,800 TFLOPS | 800 TFLOPS | 400 TFLOPS |
| FP64 DGEMM³ | 14,400 TFLOPS | 400 TFLOPS | 200 TFLOPS |
| GPU Memory \| Bandwidth | 20.7 TB HBM4 \| 1,580 TB/s | 576 GB HBM4 \| 44 TB/s | 288 GB HBM4 \| 22 TB/s |
| NVLink Bandwidth | 260 TB/s | 7.2 TB/s | 3.6 TB/s |
| NVLink-C2C Bandwidth | 65 TB/s | 1.8 TB/s | - |
| CPU Core Count | 3,168 custom NVIDIA Olympus cores (Arm® compatible) | 88 custom NVIDIA Olympus cores (Arm compatible) | - |
| CPU Memory | 54 TB LPDDR5X | 1.5 TB LPDDR5X | - |
| Total NVIDIA + HBM4 Chips | 1,296 | 30 | 12 |

1. Preliminary information. All values are up to and subject to change.
2. Dense specification.
3. Peak performance using Tensor Core-based emulation algorithms.

## 原文链接
https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72
