---
格式版本: 2
标题: "H100, L4 and Orin Raise the Bar for Inference in MLPerf | NVIDIA Blogs"
原文链接: "https://blogs.nvidia.com/blog/inference-mlperf-ai/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期位于文章标题下方的元数据区域，且带有 nvidia-article-date 类名，符合文章发布时间的特征。"
发布时间校准置信度: "1"
发布时间候选数量: 34
发布时间严格候选数量: 10
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T14:41:17+08:00"
发现时间: "2026-07-20T09:27:14+08:00"
入库时间: "2026-07-20T06:49:04.250Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=H3C"
匹配关键词:
  - "GPU"
  - "Nvlink"
相关厂家:
  - "H3C"
  - "NVIDIA"
  - "Microsoft"
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 42
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "文章主要讨论H100、L4和Orin在MLPerf推理基准测试中的性能表现，属于GPU单卡/系统级AI推理性能评测，未涉及超节点、AI Rack、机柜级系统架构、供电、散热、高速互连等核心主题。虽为NVIDIA官方发布，但内容偏向基准测试…"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T14:49:04+08:00"
AI主题相关性: 8
AI来源权威性: 14
AI新颖性: 5
AI技术细节: 5
AI商业部署信号: 5
AI完整性: 5
图片摘要:
  - "★ ./assets/img-bf721c4e.jpg | chart | H100 GPU在MLPerf推理测试中相比A100的性能优势，部分模型性能提升近4倍。"
  - "★ ./assets/img-7ca60fc3.jpg | chart | L4 GPU在MLPerf推理测试中相比T4的性能提升，最高达3.1倍。"
  - "★ ./assets/img-d37be962.jpg | chart | Jetson AGX Orin在MLPerf测试中一年内性能提升最高81%，能效提升最高63%。"
  - "✗ ./assets/img-83d2f103.jpg | photo | 通用数据中心渲染图，无具体数据或架构信息，仅作为装饰性配图。"
  - "✗ ./assets/img-1f6575f1.png | diagram | 展示Vera CPU与GPU连接，与正文MLPerf推理测试主题无关。"
  - "✗ ./assets/img-3511966c.jpg | photo | NVIDIA公司大楼及Logo照片，属于品牌宣传，与正文技术内容无关。"
  - "✗ ./assets/img-e47e99dc.png | infographic | 抽象的绿色立方体示意图，未展示具体的软件栈架构或数据，信息密度低。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/inference-mlperf-ai"
---

[MLPerf](https://www.nvidia.com/en-us/data-center/resources/mlperf-benchmarks/) remains the definitive measurement for AI performance as an independent, third-party benchmark. NVIDIA’s AI platform has consistently shown leadership across both training and inference since the inception of MLPerf, including the MLPerf Inference 3.0 benchmarks released today.

“Three years ago when we introduced A100, the AI world was dominated by computer vision. Generative AI has arrived,” said NVIDIA founder and CEO Jensen Huang.

“This is exactly why we built Hopper, specifically optimized for GPT with the Transformer Engine. Today’s MLPerf 3.0 highlights Hopper delivering 4x more performance than A100.

“The next level of Generative AI requires new AI infrastructure to train large language models with great energy efficiency. Customers are ramping Hopper at scale, building AI infrastructure with tens of thousands of Hopper GPUs connected by NVIDIA NVLink and InfiniBand.

“The industry is working hard on new advances in safe and trustworthy Generative AI. Hopper is enabling this essential work,” he said.

The latest [MLPerf results](https://www.nvidia.com/en-us/data-center/resources/mlperf-benchmarks/) show NVIDIA taking AI inference to new levels of performance and efficiency from the cloud to the edge.

Specifically, [NVIDIA H100 Tensor Core GPUs](https://www.nvidia.com/en-us/data-center/h100/) running in [DGX H100 systems](https://www.nvidia.com/en-us/data-center/dgx-h100/) delivered the highest performance in every test of AI inference, the job of running neural networks in production. Thanks to [software optimizations](https://developer.nvidia.com/blog/?p=62958&preview=1&_ppp=80347d2303), the GPUs delivered up to 54% performance gains from [their debut](https://blogs.nvidia.com/blog/hopper-mlperf-inference/) in September.

In healthcare, H100 GPUs delivered a 31% performance increase since September on 3D-UNet, the MLPerf benchmark for medical imaging.

![H100 GPU AI inference performance on MLPerf workloads](./assets/img-bf721c4e.jpg)

Powered by its [Transformer Engine](https://blogs.nvidia.com/blog/h100-transformer-engine/), the H100 GPU, based on the Hopper architecture, excelled on BERT, a transformer-based [large language model](https://blogs.nvidia.com/blog/what-are-large-language-models-used-for/) that paved the way for today’s broad use of generative AI.

Generative AI lets users quickly create text, images, 3D models and more. It’s a capability companies from startups to cloud service providers are rapidly adopting to enable new business models and accelerate existing ones.

Hundreds of millions of people are now using generative AI tools like ChatGPT — also a transformer model — expecting instant responses.

At this iPhone moment of AI, performance on inference is vital. Deep learning is now being deployed nearly everywhere, driving an insatiable need for inference performance from factory floors to online [recommendation systems](https://blogs.nvidia.com/blog/whats-a-recommender-system/).

## L4 GPUs Speed Out of the Gate

[NVIDIA L4 Tensor Core GPUs](https://www.nvidia.com/en-us/data-center/l4/) made their debut in the MLPerf tests at over 3x the speed of prior-generation T4 GPUs. Packaged in a low-profile form factor, these accelerators are designed to deliver high throughput and low latency in almost any server.

L4 GPUs ran all MLPerf workloads. Thanks to their support for the key FP8 format, their results were particularly stunning on the performance-hungry BERT model.

![NVIDIA L4 GPU AI inference performance on MLPerf workloads](./assets/img-7ca60fc3.jpg)

In addition to stellar AI performance, L4 GPUs deliver up to 10x faster image decode, up to 3.2x faster video processing and over 4x faster graphics and real-time rendering performance.

Announced two weeks ago at [GTC](https://www.nvidia.com/gtc/keynote/), these accelerators are already available from major systems makers and [cloud service providers](https://nvidianews.nvidia.com/news/nvidia-and-google-cloud-deliver-powerful-new-generative-ai-platform-built-on-the-new-l4-gpu-and-vertex-ai). L4 GPUs are the latest addition to NVIDIA’s portfolio of [AI inference platforms](https://nvidianews.nvidia.com/news/nvidia-launches-inference-platforms-for-large-language-models-and-generative-ai-workloads) launched at GTC.

## Software, Networks Shine in System Test

NVIDIA’s full-stack AI platform showed its leadership in a new MLPerf test.

The so-called network-division benchmark streams data to a remote inference server. It reflects the popular scenario of enterprise users running AI jobs in the cloud with data stored behind corporate firewalls.

On BERT, remote [NVIDIA DGX A100](https://www.nvidia.com/en-us/data-center/dgx-a100/) systems delivered up to 96% of their maximum local performance, slowed in part because they needed to wait for CPUs to complete some tasks. On the ResNet-50 test for computer vision, handled solely by GPUs, they hit the full 100%.

Both results are thanks, in large part, to [NVIDIA Quantum Infiniband](https://www.nvidia.com/en-au/networking/infiniband/qm8700/) networking, [NVIDIA ConnectX SmartNICs](https://www.nvidia.com/en-us/networking/ethernet-adapters/) and software such as [NVIDIA GPUDirect](https://developer.nvidia.com/gpudirect).

## Orin Shows 3.2x Gains at the Edge

Separately, the NVIDIA Jetson AGX Orin system-on-module delivered gains of up to 63% in energy efficiency and 81% in performance compared with its results a year ago. Jetson AGX Orin supplies inference when AI is needed in confined spaces at low power levels, including on systems powered by batteries.

![Jetson AGX Orin AI inference performance on MLPerf benchmarks](./assets/img-d37be962.jpg)

For applications needing even smaller modules drawing less power, the Jetson Orin NX 16G shined in its debut in the benchmarks. It delivered up to 3.2x the performance of the prior-generation Jetson Xavier NX processor.

## A Broad NVIDIA AI Ecosystem

The MLPerf results show NVIDIA AI is backed by the industry’s broadest ecosystem in machine learning.

Ten companies submitted results on the NVIDIA platform in this round. They came from the Microsoft Azure cloud service and system makers including ASUS, [Dell Technologies](https://infohub.delltechnologies.com/p/dell-servers-excel-in-mlcommonstm-inference-3-0-performance/), GIGABYTE, New H3C Information Technologies, [Lenovo](https://lenovopress.lenovo.com/lp1714-exceeding-the-expectations-of-ai-performance-through-mlperf), Nettrix, Supermicro and xFusion.

Their work shows users can get great performance with NVIDIA AI both in the cloud and in servers running in their own data centers.

NVIDIA partners participate in MLPerf because they know it’s a valuable tool for customers evaluating AI platforms and vendors. Results in the latest round demonstrate that the performance they deliver today will grow with the NVIDIA platform.

## Users Need Versatile Performance

NVIDIA AI is the only platform to run all MLPerf inference workloads and scenarios in data center and edge computing. Its versatile performance and efficiency make users the real winners.

Real-world applications typically employ many neural networks of different kinds that often need to deliver answers in real time.

For example, an AI application may need to understand a user’s spoken request, classify an image, make a recommendation and then deliver a response as a spoken message in a human-sounding voice. Each step requires a different type of AI model.

The MLPerf benchmarks cover these and other popular AI workloads. That’s why the tests ensure IT decision makers will get performance that’s dependable and flexible to deploy.

Users can rely on MLPerf results to make informed buying decisions, because the tests are transparent and objective. The benchmarks enjoy backing from a broad group that includes Arm, Baidu, Facebook AI, Google, Harvard, Intel, Microsoft, Stanford and the University of Toronto.

## Software You Can Use

The software layer of the NVIDIA AI platform, [NVIDIA AI Enterprise](https://www.nvidia.com/en-us/data-center/products/ai-enterprise/), ensures users get optimized performance from their infrastructure investments as well as the enterprise-grade support, security and reliability required to run AI in the corporate data center.

All the software used for these tests is available from [the MLPerf repository](https://github.com/mlcommons/inference_results_v3.0), so anyone can get these world-class results.

Optimizations are continuously folded into containers available on [NGC](https://ngc.nvidia.com/catalog), NVIDIA’s catalog for GPU-accelerated software. The catalog hosts [NVIDIA TensorRT](https://developer.nvidia.com/tensorrt), used by every submission in this round to optimize AI inference.

Read this [technical blog](https://developer.nvidia.com/blog/?p=62958&preview=1&_ppp=80347d2303) for a deeper dive into the optimizations fueling NVIDIA’s MLPerf performance and efficiency.

![H100 GPU AI inference performance on MLPerf workloads](./assets/img-bf721c4e.jpg)

![NVIDIA L4 GPU AI inference performance on MLPerf workloads](./assets/img-7ca60fc3.jpg)

![Jetson AGX Orin AI inference performance on MLPerf benchmarks](./assets/img-d37be962.jpg)
