---
格式版本: 2
标题: "NVIDIA, Partners Deliver Top AI Training Results on MLPerf | NVIDIA Blogs"
原文链接: "https://blogs.nvidia.com/blog/mlperf-ai-training-partners/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期来自HTML metadata中的article-date字段，位于标题附近，符合文章发布时间的特征。"
发布时间校准置信度: "100"
发布时间候选数量: 41
发布时间严格候选数量: 11
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T11:29:37+08:00"
发现时间: "2026-07-20T09:24:55+08:00"
入库时间: "2026-07-20T03:37:59.790Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=PCIe%20switch"
匹配关键词:
  - "PCIe switch"
  - "GPU"
相关厂家:
  - "NVIDIA"
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 48
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "资料为NVIDIA官方博客，但内容聚焦于MLPerf基准测试中A100 GPU的训练性能表现，未涉及超节点、AI Rack、机柜级系统架构、供电、散热、高速互连等核心主题。技术细节主要围绕CUDA Graphs和SHARP软件优化，缺乏硬…"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T11:37:59+08:00"
AI主题相关性: 8
AI来源权威性: 14
AI新颖性: 4
AI技术细节: 6
AI商业部署信号: 8
AI完整性: 8
图片摘要:
  - "★ ./assets/img-a0e3f03f.jpg | chart | MLPerf基准测试显示NVIDIA DGX SuperPOD在8个模型训练中速度最快，大幅领先Graphcore、Habana和Intel等竞品。"
  - "★ ./assets/img-afbb0557.jpg | chart | MLPerf测试表明NVIDIA A100单芯片性能最强，在全部8个AI模型测试中均创下纪录，性能远超竞争对手。"
  - "★ ./assets/img-93680aab.jpg | chart | 图表展示从V100到A100及软件迭代带来的MLPerf性能显著提升，部分模型提升超6倍，印证整体性能增长。"
  - "✗ ./assets/img-83d2f103.jpg | other | 图片标题为另一篇文章主题（能效），与正文MLPerf训练结果无直接关联，且为通用机房渲染图。"
  - "✗ ./assets/img-1f6575f1.png | diagram | 图片标题涉及Vera CPU，属于另一篇文章内容，与正文A100 GPU训练主题无关。"
  - "✗ ./assets/img-3511966c.jpg | photo | 图片为NVIDIA大楼实拍及Logo，标题为另一篇文章（基础设施构建），属于品牌宣传/无关配图。"
  - "✗ ./assets/img-e47e99dc.png | other | 图片标题涉及推理（Inference）软件栈，与正文训练（Training）主题不符，属于另一篇文章配图。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/mlperf-ai-training-partners"
---

NVIDIA’s partners are delivering GPU-accelerated systems that train AI models faster than anyone on the planet, according to the [latest MLPerf results](https://www.nvidia.com/en-us/data-center/mlperf/) released today.

Seven companies put at least a dozen commercially available systems, the majority [NVIDIA-Certified](https://www.nvidia.com/en-us/data-center/products/certified-systems/), to the test in the industry benchmarks. Dell, Fujitsu, GIGABYTE, [Inspur Electronic Information](https://www.inspursystems.com/newsroom/inspur-ai-servers-lead-mlperf-training-v1-benchmarks/), [Lenovo](https://www.lenovoxperience.com/newsDetail/283yi044hzgcdv7snkrmmx9ozw87e6voioc5v5wxfx3jzo6a), [Nettrix](https://www.nettrix.com.cn/cut?id=1084) and Supermicro joined NVIDIA to demonstrate industry-leading results training neural networks with [NVIDIA A100 Tensor Core GPUs](https://www.nvidia.com/en-us/data-center/a100/).

Only NVIDIA and its partners ran all eight workloads in the latest round of benchmarks. Together, our work made up more than three-quarters of all submissions, and the results were stunning.

Compared to last year’s scores, we delivered up to 3.5x more performance. For massive jobs that demand the most muscle, we mustered resources from a record 4,096 GPUs, more than any other submission.

## Why MLPerf Matters

This is the fourth and strongest showing for the NVIDIA ecosystem in training tests from [MLPerf](https://www.nvidia.com/en-us/data-center/mlperf/), an industry benchmarking group formed in May 2018.

MLPerf gives users the confidence to make informed buying decisions. It’s backed by dozens of industry leaders including Alibaba, Arm, Baidu, Google, Intel and NVIDIA, so the tests are transparent and objective.

The benchmarks are based on today’s most popular AI workloads and scenarios, covering computer vision, natural-language processing, recommendation systems, reinforcement learning and more. And the training benchmarks focus on what users care about most — time to train a new AI model.

## Speed + Flexibility = Productivity

Ultimately, the return on a customer’s infrastructure investment depends on their productivity. That comes from an ability to be both fast and flexible when running these many kinds of AI workloads.

That’s why users need flexible, yet powerful systems that can get a variety of AI models into production fast, speeding time to market and maximizing the productivity of their valuable data science teams.

In the latest MLPerf results, the NVIDIA AI platform set performance records by training models in the shortest time across all eight benchmarks in the commercially available submissions category.

![MLPerf at scale results](./assets/img-a0e3f03f.jpg)

Selene, based on an NVIDIA DGX SuperPOD, set all eight records on commercially available systems.

We ran the at-scale tests on [Selene](https://blogs.nvidia.com/blog/making-selene-pandemic-ai/), the fastest commercial AI supercomputer in the world, according to the [latest TOP500 rankings](https://blogs.nvidia.com/blog/top500-ai-cloud-native/). It’s based on the same [NVIDIA DGX SuperPOD](https://www.nvidia.com/en-us/data-center/dgx-superpod/) architecture that powers a dozen other systems on the list.

The ability to scale to large clusters is the toughest challenge in AI and one of our core strengths.

In chip-to-chip comparisons, NVIDIA and its partners set records across all eight benchmarks in the latest tests on commercially available systems.

![MLPerf chip to chip results](./assets/img-afbb0557.jpg)

A100 GPUs set all eight records in the category of commercially available systems.

Overall, the results below show our performance rose up to 6.5x in 2.5 years, a testament to work across the full-stack NVIDIA platform of GPUs, systems and software.

![MLPerf improvements](./assets/img-93680aab.jpg)

NVIDIA AI delivered continuous gains with full-stack improvements.

## Broad Ecosystem Offers Best Value, Choice

The MLPerf results demonstrate performance across a variety of NVIDIA-based AI platforms with plenty of new and innovative systems. They span entry-level edge servers to AI supercomputers that accommodate thousands of GPUs.

The seven partners participating in the latest benchmarks are among nearly two dozen cloud-service providers and OEMs with products or plans for online instances, servers and PCIe cards using NVIDIA A100 GPUs, including nearly 40 [NVIDIA-Certified Systems](https://www.nvidia.com/en-us/data-center/products/certified-systems/).

Our ecosystem offers customers choices in a wide range of deployment models — from instances that are rentable by the minute to on-prem servers and managed services — providing the most value per dollar in the industry.

Results across all the MLPerf tests show our performance keeps rising over time. That comes from a platform with software that’s mature and constantly improving, so teams can get started fast with systems that keep getting better.

## How We Did It

It’s the second round of MLPerf tests for our A100 GPUs. Speedups came from advances detailed in [a separate article](https://developer.nvidia.com/blog/mlperf-v1-0-training-benchmarks-insights-into-a-record-setting-performance/) that spanned GPUs, systems, networking and AI software.

For example, our engineers found a way to launch full neural network models using [CUDA Graphs](https://developer.nvidia.com/blog/cuda-graphs/), a software package of NVIDIA CUDA operations and their dependencies. That eliminated CPU bottlenecks in past tests that released AI models as a chain of many individual components, called kernels.

In addition, tests at scale used [NVIDIA SHARP](https://docs.mellanox.com/display/SHARPv200), software that consolidates multiple communications jobs inside a network switch, reducing network traffic and time waiting for a CPU.

The combination of the CUDA Graphs and SHARP allowed training jobs in the data center to access a record number of GPUs. It’s the muscle needed in many areas such as natural-language processing, where AI models are growing to include billions of parameters.

Other gains came from expanded memory on the latest A100 GPUs that increase memory bandwidth nearly 30 percent to more than 2 TB/s.

## Customers Care About MLPerf

A wide variety of AI users find these benchmarks useful.

“The MLPerf benchmark provides a transparent apples-to-apples comparison across multiple AI platforms to showcase actual performance in diverse real-world use cases,” said a spokesman for Sweden’s Chalmers University, which conducts research across areas from nanotechnology to climate studies.

The benchmarks help users find the AI products that meet the requirements of some of the world’s largest and most advanced factories. For example, TSMC, a global leader in chip manufacturing, uses machine learning to improve optical proximity correction (OPC) and etch simulation.

“To fully realize the potential of machine learning in model training and inference, we’re working with the NVIDIA engineering team to port our Maxwell simulation and inverse lithography technology engine to GPUs and see very significant speedups. The MLPerf benchmark is an important factor in our decision making,” said Danping Peng, director of TSMC’s OPC department.

## Traction in Medicine, Manufacturing

The benchmarks also are useful for researchers pushing the limits of AI to improve healthcare.

“We’ve worked closely with NVIDIA to bring innovations like 3DUNet to the healthcare market. Industry-standard MLPerf benchmarks provide relevant performance data to the benefit of IT organizations and developers to get the right solution to accelerate their specific projects and applications,” said Klaus Maier-Hein, head of medical image computing at DKFZ, the German cancer research center.

A world leader in research and manufacturing, Samsung also uses MLPerf benchmarks as it implements AI to boost product performance and manufacturing productivity.

“Productizing these AI advances requires us to have the best computing platform available. The MLPerf benchmark streamlines our selection process by providing us with an open, direct evaluation method to assess uniformly across platform vendors,” said a spokesperson for Samsung Electronics.

## Get These Same Results, Tools

All the software we used for the latest submissions is available from the MLPerf repository, so anyone can reproduce our benchmark results. We continually add this code into our deep learning frameworks and containers available on [NGC](https://ngc.nvidia.com/), our software hub for GPU applications.

It’s part of a full-stack AI platform, proven in the latest industry benchmarks, and available from a variety of partners to tackle real AI jobs today.

![MLPerf at scale results](./assets/img-a0e3f03f.jpg) Selene, based on an NVIDIA DGX SuperPOD, set all eight records on commercially available systems.

![MLPerf chip to chip results](./assets/img-afbb0557.jpg) A100 GPUs set all eight records in the category of commercially available systems.

![MLPerf improvements](./assets/img-93680aab.jpg) NVIDIA AI delivered continuous gains with full-stack improvements.
