---
格式版本: 2
标题: "Microsoft Azure Unveils World’s First NVIDIA GB300 NVL72 Supercomputing Cluster for OpenAI | NVIDIA Blog"
原文链接: "https://blogs.nvidia.com/blog/microsoft-azure-worlds-first-gb300-nvl72-supercomputing-cluster-openai/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期来自HTML metadata中的article-date字段，位于标题附近，符合文章发布时间的特征。"
发布时间校准置信度: "1"
发布时间候选数量: 36
发布时间严格候选数量: 12
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T10:25:01+08:00"
发现时间: "2026-07-20T09:24:00+08:00"
入库时间: "2026-07-20T02:28:25.798Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=Scale-up"
匹配关键词:
  - "Scale-up"
  - "NVL72"
  - "GPU"
  - "Liquid Cooling"
  - "Nvlink"
相关厂家:
  - "NVIDIA"
  - "Microsoft"
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "是"
AI打分: 92
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "NVIDIA官方发布，直接讨论GB300 NVL72超节点/机柜级AI系统，含详细架构、NVLink/InfiniBand互连、液冷、内存规格及OpenAI/Microsoft大规模部署信号，信息完整权威。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T10:28:25+08:00"
AI主题相关性: 20
AI来源权威性: 15
AI新颖性: 18
AI技术细节: 19
AI商业部署信号: 15
AI完整性: 5
图片摘要:
  - "★ ./assets/img-83d2f103.jpg | diagram | 展示 NVIDIA GB300 NVL72 液冷机架及其在 Azure 超算集群中的大规模互连部署形态。"
  - "✗ ./assets/img-1f6575f1.png | diagram | 图片主题为 NVIDIA Vera CPU，而正文主要介绍基于 Grace CPU 的 GB300 系统，产品代次不符。"
  - "✗ ./assets/img-3511966c.jpg | photo | NVIDIA 公司 Logo 及建筑实拍，属于品牌宣传/装饰性配图，无具体技术信息。"
  - "✗ ./assets/img-e47e99dc.png | infographic | 文件名暗示与 AWS 相关或为通用软件栈图，与正文 Azure GB300 硬件集群主题不直接相关。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/microsoft-azure-worlds-first-gb300-nvl72-supercomputing-cluster-openai"
---

Microsoft Azure today [announced](https://azure.microsoft.com/en-us/blog/?p=47052) the new NDv6 GB300 VM series, delivering the industry’s first supercomputing-scale production cluster of [NVIDIA GB300 NVL72](https://www.nvidia.com/en-us/data-center/gb300-nvl72/) systems, purpose-built for OpenAI’s most demanding AI inference workloads.

This supercomputer-scale cluster features over 4,600 NVIDIA Blackwell Ultra GPUs connected via the [NVIDIA Quantum-X800 InfiniBand](https://www.nvidia.com/en-us/networking/products/infiniband/quantum-x800/) networking platform. Microsoft’s unique systems approach applied radical engineering to memory and networking to provide the massive scale of compute required to achieve high inference and training throughput for reasoning models and agentic AI systems.

Today’s achievement is the result of years of deep partnership between NVIDIA and Microsoft purpose-building AI infrastructure for the world’s most demanding AI workloads and to deliver infrastructure for the next frontier of AI. It marks another leadership moment, ensuring that leading-edge AI drives innovation in the United States.

“Delivering the industry’s first at-scale NVIDIA GB300 NVL72 production cluster for frontier AI is an achievement that goes beyond powerful silicon — it reflects Microsoft Azure and NVIDIA’s shared commitment to optimize all parts of the modern AI data center,” said Nidhi Chappell, corporate vice president of Microsoft Azure AI Infrastructure.

“Our collaboration helps ensure customers like OpenAI can deploy next-generation infrastructure at unprecedented scale and speed.”

## Inside the Engine: The NVIDIA GB300 NVL72

At the heart of Azure’s new NDv6 GB300 VM series is the liquid-cooled, rack-scale NVIDIA GB300 NVL72 system. Each rack is a powerhouse, integrating 72 NVIDIA Blackwell Ultra GPUs and 36 NVIDIA Grace CPUs into a single, cohesive unit to accelerate training and inference for massive AI models.

The system provides a staggering 37 terabytes of fast memory and 1.44 exaflops of FP4 Tensor Core performance per VM, creating a massive, unified memory space essential for reasoning models, agentic AI systems and complex multimodal generative AI.

NVIDIA Blackwell Ultra is supported by the full-stack NVIDIA AI platform, including collective communication libraries that tap into new formats like [NVFP4](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/) for breakthrough training performance, as well as compiler technologies like [NVIDIA Dynamo](https://www.nvidia.com/en-us/ai/dynamo/) for the highest inference performance in reasoning AI.

The NVIDIA Blackwell Ultra platform excels at both training and inference. In the recent [MLPerf Inference v5.1](https://blogs.nvidia.com/blog/mlperf-inference-blackwell-ultra/) benchmarks, NVIDIA GB300 NVL72 systems delivered record-setting performance using NVFP4. Results included up to 5x higher throughput per GPU on the 671-billion-parameter DeepSeek-R1 reasoning model compared with the NVIDIA Hopper architecture, along with leadership performance on all newly introduced benchmarks like the Llama 3.1 405B model.

## The Fabric of a Supercomputer: NVLink Switch and NVIDIA Quantum-X800 InfiniBand

To connect over 4,600 Blackwell Ultra GPUs into a single, cohesive supercomputer, Microsoft Azure’s cluster relies on a two-tiered NVIDIA networking architecture designed for both scale-up performance within the rack and scale-out performance across the entire cluster.

Within each GB300 NVL72 rack, the fifth-generation [NVIDIA NVLink Switch](https://www.nvidia.com/en-us/data-center/nvlink/) fabric provides 130 TB/s of direct, all-to-all bandwidth between the 72 Blackwell Ultra GPUs. This transforms the entire rack into a single, unified accelerator with a shared memory pool — a critical design for massive, memory-intensive models.

To scale beyond the rack, the cluster uses the NVIDIA Quantum-X800 InfiniBand platform, purpose-built for trillion-parameter-scale AI. Featuring NVIDIA ConnectX-8 SuperNICs and Quantum-X800 switches, NVIDIA Quantum-X800 provides 800 Gb/s of bandwidth per GPU, ensuring seamless communication across all 4,608 GPUs.

Microsoft Azure’s cluster also uses NVIDIA Quantum-X800’s advanced adaptive routing, telemetry-based congestion control and performance isolation capabilities, as well as NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) v4, which accelerates operations to significantly boost the efficiency of large-scale training and inference.

## Driving the Future of AI

Delivering the world’s first production NVIDIA GB300 NVL72 cluster at this scale required a reimagination of every layer of Microsoft’s data center — from custom liquid cooling and power distribution to a reengineered software stack for orchestration and storage.

This latest milestone marks a big step forward in building the infrastructure that will unlock the future of AI. As Azure scales to its goal of deploying hundreds of thousands of NVIDIA Blackwell Ultra GPUs, even more innovations are poised to emerge from customers like OpenAI.

*Learn more about this announcement on the* [*Microsoft Azure blog*](https://azure.microsoft.com/en-us/blog/?p=47052)*.*

![Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency](./assets/img-83d2f103.jpg)
