---
格式版本: 2
标题: "Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency | NVIDIA Blog"
原文链接: "https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/"
发布日期: "2026-07-14"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-07-14T08:00:20-07:00: Jul 14, 2026"
发布时间校准原因: "该日期来自HTML metadata中的article-date字段，且位于标题附近，符合文章发布时间的特征。"
发布时间校准置信度: "100"
发布时间候选数量: 36
发布时间严格候选数量: 12
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T10:24:54+08:00"
发现时间: "2026-07-20T09:24:00+08:00"
入库时间: "2026-07-20T02:35:05.840Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=Scale-up"
匹配关键词:
  - "Scale-up"
  - "GPU"
  - "Liquid Cooling"
  - "NVL72"
  - "Nvlink"
  - "Vera Rubin"
相关厂家:
  - "NVIDIA"
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "是"
AI打分: 92
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "NVIDIA官方发布，深度讨论AI基础设施能效与机柜级系统（NVL72/GB300/Vera Rubin），涵盖Scale-up互连、NVLink Switch、液冷、功耗调度及量产部署信号，技术细节丰富，来源权威，内容完整。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T10:35:05+08:00"
AI主题相关性: 19
AI来源权威性: 15
AI新颖性: 18
AI技术细节: 18
AI商业部署信号: 12
AI完整性: 10
图片摘要:
  - "★ ./assets/img-f352e1f2.png | chart | Blackwell GB300对比Hopper在DeepSeek-v4-PRO模型下实现25倍兆瓦吞吐量提升的性能曲线。"
  - "★ ./assets/img-e43dc157.png | chart | Blackwell GB300对比Hopper在GLM 5.1模型下实现20倍兆瓦吞吐量提升的性能曲线。"
  - "★ ./assets/img-8987ef77.png | chart | Blackwell GB300对比Hopper在Kimi系列模型下实现10倍兆瓦吞吐量提升的性能曲线。"
  - "✓ ./assets/img-83d2f103.jpg | diagram | 展示大规模AI工厂基础设施布局的概念渲染图，体现高密度机柜部署。"
  - "✓ ./assets/img-1f6575f1.png | diagram | 展示NVIDIA Vera CPU与GPU协同处理Agentic AI任务的简化互连示意图。"
  - "✗ ./assets/img-3511966c.jpg | photo | 品牌宣传/装饰性配图，与正文技术内容无直接关联。"
  - "✓ ./assets/img-e47e99dc.png | photo | NVIDIA Blackwell芯片或服务器组件的产品渲染图，展示硬件外观。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency"
---

Power is AI infrastructure’s inescapable constraint. How many [tokens](https://blogs.nvidia.com/blog/ai-tokens-explained/) an AI factory can generate within a fixed power budget determines its revenue and profitability. Because of this, performance per watt — a metric that can’t be gamed, only earned through real-world results — is the foundation for AI factories.

As agentic AI drives token demand higher, the infrastructure decisions organizations make today will determine who scales and who doesn’t in a power-constrained world.

Virtually every frontier AI model today runs on a [mixture-of-experts](https://blogs.nvidia.com/blog/mixture-of-experts-frontier-models/) (MoE) architecture. Serving these large-scale models efficiently means GPU domain size — the number of GPUs connected over an ultrafast, scale-up interconnect — matters, and bigger is better.

While the NVIDIA Hopper generation set the standard with an eight-GPU domain, the scale of frontier AI today has outgrown it. Serving MoE with a 72-GPU domain demands full-stack codesign and the operational depth earned from running these models under real production load.

With the [NVIDIA Blackwell NVL72 platform](https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/), that foundation is already built and proven, delivering the highest performance per watt to maximize revenues and the lowest token cost to maximize profit margins. It’s this foundation that the [NVIDIA Vera Rubin](https://www.nvidia.com/en-us/data-center/technologies/rubin/) platform builds upon next to further elevate rack-scale energy efficiency.

## Maximizing Performance per Watt for Frontier AI

Each new generation of frontier models brings architectural changes that unlock greater intelligence while demanding new optimizations to run efficiently at scale.

Across the newest generation of leading open models, NVIDIA GB300 NVL72 delivers up to 25x performance per watt compared with the NVIDIA Hopper generation — showcasing that MoE performance improves when moving from an 8-GPU to 72-GPU domain size. These numbers reflect where Blackwell stands today, a starting point that continues to improve.

Any single number only tells part of the story. Different workloads demand different operating points: some optimize for latency, others for throughput and cost — and most need to move between the two.

To best represent these operating points, NVIDIA showcases Pareto curves for each model rather than a single point and provides tools such as [DynoSim](https://developer.nvidia.com/blog/dynosim-simulating-the-pareto-frontier/) to help teams find their optimal point on the Pareto frontier before spending a single GPU-hour on validation.

![图片](./assets/img-f352e1f2.png)

NVIDIA GB300 NVL72 systems deliver up to 25x performance per watt over NVIDIA Hopper on DeepSeek V4 Pro. Source: SemiAnalysis InferenceX

![图片](./assets/img-e43dc157.png)

On GLM5.1 NVIDIA GB300 NVL72 systems deliver up to 20x performance per watt over NVIDIA Hopper. Source: SemiAnalysis InferenceX

![图片](./assets/img-8987ef77.png)

NVIDIA GB300 NVL72 systems deliver up to 10x performance per watt over NVIDIA Hopper for Kimi K2.6, a model purpose-built for long-horizon agentic tasks. Source: SemiAnalysis InferenceX

The performance per watt NVIDIA Blackwell delivers is a result of extreme codesign: every component of the rack-scale system, from silicon to [software](https://blogs.nvidia.com/blog/inference-software-lowest-token-cost/), designed together to maximize token throughput for AI inference workloads. That codesign touches every layer of the stack.

For example, [NVIDIA NVLink Switch](https://www.nvidia.com/en-us/data-center/nvlink/), critical for rack-scale performance, is purpose-built to unlock massive scale-up GPU domains, not adapted from general-purpose networking. Now in its sixth generation with the Vera Rubin platform, its capabilities are designed specifically for AI workloads such as SHARP, which performs in-network computing directly in the switch, offloading work from the GPUs themselves.

NVIDIA’s [inference software stack](https://blogs.nvidia.com/blog/inference-software-lowest-token-cost/), including NVIDIA Dynamo and TensorRT LLM, as well as SGLang and vLLM, is built to run the full range of optimizations: NVFP4 quantization, disaggregated serving, large-scale expert parallelism, KV-aware routing, KV cache offloading and more. These stack together to multiply the performance each GPU delivers. Moreover, software keeps improving performance over time: On DeepSeek V4, performance per watt improved by up to 5x in a single month.

In AI factories, power lost to cooling and rack-level inefficiencies can mean only about 60% of the electricity pulled from the grid turns into useful AI work. NVIDIA DSX MaxLPS, the power-and-efficiency software in the [NVIDIA DSX](https://www.nvidia.com/en-us/data-center/products/dsx/) platform, closes that gap by shifting power between GPUs and racks in real time, supporting warm-water liquid cooling and using techniques like power steering to wring more performance. This enables operators to run up to 40% more GPUs within the same power budget.

## Production Is Where It Counts

Rack-scale reliability at AI factory scale is hard-won. Rack-scale systems introduce failure modes that single-node deployments never encounter, and handling them requires engineering rigor and time in production.

NVIDIA Blackwell NVL72 systems continues to set the standard across a diverse range of models and production use cases delivering sustained performance, rack-level reliability and economics that hold under real traffic day after day.

That’s why leading AI labs such as Anthropic, OpenAI and SpaceXAI use NVIDIA Blackwell NVL72 systems to run inference.

In addition, a variety of inference service providers and AI natives use the Blackwell platform to deploy open models in production.

[CoreWeave has deployed Kimi K2.6](https://www.coreweave.com/blog/coreweave-is-now-the-fastest-at-inference-on-the-best-open-source-model-kimi-k2-6) on NVIDIA GB300 NVL72, combining NVFP4 quantization and EAGLE3 speculative decoding to maximize inference performance.

Perplexity runs [Qwen3 235B](https://research.perplexity.ai/articles/advancing-search-augmented-language-models) and post-trained Qwen3.5-397B-A17B on NVIDIA GB200 NVL72 for its AI agent platform, serving millions of queries daily with the latency and reliability that consumers need.

Fireworks AI deploys GLM 5.2 on the NVIDIA Blackwell platform, enabling production deployments for customers including Cursor and Factory AI.

This accumulated production experience, built across generations of frontier models and real-world deployments, is what gives NVIDIA Vera Rubin its head start.

*Learn more about the NVIDIA Vera Rubin platform in this* [*technical blog*](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/) *and find details on the* [*NVIDIA DSX AI factory-scale platform and DSX MaxLPS*](https://docs.nvidia.com/dsx)*.*

![图片](./assets/img-f352e1f2.png)NVIDIA GB300 NVL72 systems deliver up to 25x performance per watt over NVIDIA Hopper on DeepSeek V4 Pro. Source: SemiAnalysis InferenceX![图片](./assets/img-e43dc157.png)On GLM5.1 NVIDIA GB300 NVL72 systems deliver up to 20x performance per watt over NVIDIA Hopper. Source: SemiAnalysis InferenceX![图片](./assets/img-8987ef77.png)NVIDIA GB300 NVL72 systems deliver up to 10x performance per watt over NVIDIA Hopper for Kimi K2.6, a model purpose-built for long-horizon agentic tasks. Source: SemiAnalysis InferenceX

![Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency](./assets/img-83d2f103.jpg)

![AI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale Matters](./assets/img-1f6575f1.png)

![How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost](./assets/img-e47e99dc.png)
