---
格式版本: 2
标题: "GB200 NVL72 | NVIDIA"
原文链接: "https://www.nvidia.com/en-us/data-center/gb200-nvl72"
发布日期: "2026-05-29"
发现时间: "2026-05-29T10:48:33+08:00"
入库时间: "2026-05-29T03:44:46.576Z"
来源平台: "Tavily"
搜索渠道: "tavily_web"
搜索词: "NVL72"
匹配关键词:
  - "NVL72"
  - "GPU"
  - "Liquid Cooling"
  - "NVIDIA"
  - "Scale-up"
  - "Google"
相关厂家:
  - "NVIDIA"
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "XCrawl Scrape"
清洗工具: "XCrawl Markdown + LLM 正文裁剪"
AI优质: "是"
AI打分: 90
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "官方产品页，直接聚焦GB200 NVL72整机柜AI架构，明确给出72卡NVLink域、130TB/s互连带宽、液冷设计及FP4/FP8精度等核心参数，技术细节扎实，来源权威，符合超节点/机柜级AI基础设施高价值标准。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-06-03T00:04:28+08:00"
AI主题相关性: 20
AI来源权威性: 15
AI新颖性: 18
AI技术细节: 18
AI商业部署信号: 10
AI完整性: 9
采集批次: "2026年5月29日10点47分48秒"
采集批次ID: "20260529-104748-298"
去重键: "https://www.nvidia.com/en-us/data-center/gb200-nvl72"
---

# NVIDIA GB200 NVL72
Powering the era of accelerated computing. [Read Datasheet](https://nvdam.widen.net/s/wwnsxrhm2w/blackwell-datasheet-3384703)

## Overview
### Unlocking Real-Time Trillion-Parameter Models
The NVIDIA GB200 NVL72 connects 36 Grace CPUs and 72 Blackwell GPUs in a rack-scale, liquid-cooled design. It boasts a 72-GPU NVIDIA NVLink™ domain that acts as a single, massive GPU and delivers 30x faster real-time trillion-parameter large language model (LLM) inference, with 10x greater performance for [mixture-of-experts (MoE) architectures](https://blogs.nvidia.com/blog/mixture-of-experts-frontier-models/). The GB200 Grace Blackwell Superchip is a key component of the [NVIDIA GB200 NVL72](https://developer.nvidia.com/blog/nvidia-gb200-nvl72-delivers-trillion-parameter-llm-training-and-real-time-inference/), connecting two high-performance NVIDIA Blackwell Tensor Core GPUs and an NVIDIA Grace™ CPU using the NVLink-C2C interconnect to the two Blackwell GPUs.
### The Blackwell Rack-Scale Architecture for Real-Time Trillion-Parameter Inference and Training
The NVIDIA GB200 NVL72 is an exascale computer in a single rack. With 72 NVIDIA Blackwell GPUs interconnected by the largest NVIDIA NVLink domain ever offered, NVLink Switch System provides 130 terabytes per second (TB/s) of low-latency GPU communications for AI and high-performance computing (HPC) workloads. [Tech Blog](https://developer.nvidia.com/blog/nvidia-gb200-nvl72-delivers-trillion-parameter-llm-training-and-real-time-inference/)

## Highlights
### Supercharging Next-Generation AI and Accelerated Computing
### LLM Inference 30x vs. NVIDIA H100 GPU
### LLM Training 4x vs. H100
### Energy Efficiency 25xvs. H100
### Data Processing 18x vs. CPU
LLM inference and energy efficiency: TTL = 50 milliseconds (ms) real time, FTL = 5s, 32,768 input/1,024 output, NVIDIA HGX™ H100 scaled over InfiniBand (IB) vs. GB200 NVL72, training 1.8T MOE 4096x HGX H100 scaled over IB vs. 456x GB200 NVL72 scaled over IB. Cluster size: 32,768 A database join and aggregation workload with Snappy / Deflate compression derived from TPC-H Q4 query. Custom query implementations for x86, H100 single GPU and single GPU from GB200 NLV72 vs. Intel Xeon 8480+ Projected performance subject to change.
### Real-Time LLM Inference
GB200 NVL72 introduces cutting-edge capabilities and a second-generation Transformer Engine, which enables FP4 AI. When coupled with fifth-generation NVIDIA NVLink, it delivers 30x faster real-time LLM inference performance for trillion-parameter language models. This advancement is made possible with a new generation of Tensor Cores, which introduce new microscaling formats, optimized for high-throughput, low-latency AI Inference. Additionally, the GB200 NVL72 uses NVLink and liquid cooling to create a single massive 72-GPU rack that can overcome communication bottlenecks.
### Massive-Scale Training
GB200 NVL72 features a faster second-generation Transformer Engine, offering FP8 precision and enabling a remarkable 4x faster training for large language models at scale. This breakthrough is complemented by the fifth-generation NVLink, which provides 1.8 TB/s of GPU-to-GPU interconnect, InfiniBand networking, and NVIDIA Magnum IO™ software.
### Energy-Efficient Infrastructure
Liquid-cooled GB200 NVL72 racks reduce a data center’s carbon footprint and energy consumption. Liquid cooling increases compute density, reduces the amount of floor space used, and facilitates high-bandwidth, low-latency GPU communication with large [NVLink domain architectures](https://www.nvidia.com/en-us/data-center/nvlink/). Compared to NVIDIA H100 air-cooled infrastructure, GB200 delivers 25x more performance at the same power, while reducing water consumption.
### Data Processing
Databases play critical roles in handling, processing, and analyzing large volumes of data for enterprises. GB200 takes advantage of the high-bandwidth memory performance, [NVLink-C2C](https://www.nvidia.com/en-us/data-center/nvlink-c2c/), and dedicated decompression engines in the [NVIDIA Blackwell architecture](https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/) to speed up key database queries by 18x compared to CPU and deliver a 5x better TCO.

## NVIDIA GB200 NVL4
NVIDIA GB200 NVL4 unlocks the future of converged HPC and AI, delivering revolutionary performance through a bridge connecting four NVIDIA NVLink Blackwell GPUs unified with two Grace CPUs over NVLink-C2C interconnect. Compatible with liquid-cooled NVIDIA MGX™ modular servers, it provides up to 2x performance for scientific computing, AI for science training, and inference applications over the prior generation. [Read Datasheet](https://nvdam.widen.net/s/wwnsxrhm2w/blackwell-datasheet-3384703)

## Features
### Technological Breakthroughs
### Blackwell Architecture
The NVIDIA Blackwell architecture delivers groundbreaking advancements in accelerated computing, powering a new era of computing with unparalleled performance, efficiency, and scale. [Learn More](https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/)
### NVIDIA Grace CPU
The NVIDIA Grace CPU is a breakthrough processor designed for modern data centers running AI, cloud, and HPC applications. It provides outstanding performance and memory bandwidth with 2x the energy efficiency of today’s leading server processors. [Learn More](https://www.nvidia.com/en-us/data-center/grace-cpu-superchip/)
### Fifth-Generation NVIDIA NVLink
Unlocking the full potential of exascale computing and trillion-parameter AI models requires swift, seamless communication between every GPU in a server cluster. The fifth generation of NVLink is a scale–up interconnect that unleashes accelerated performance for trillion- and multi-trillion-parameter AI models. [Learn About NVLink and NVLink Switch](https://www.nvidia.com/e…
