---
格式版本: 2
标题: "NVIDIA DSX OS Delivers Open, Modular Software for Operating AI Factories at Scale | NVIDIA Technical Blog"
原文链接: "https://developer.nvidia.com/blog/nvidia-dsx-os-delivers-open-modular-software-for-operating-ai-factories-at-scale/"
发布日期: "2026-05-31"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "div class=post-info: May 31, 2026"
发布时间校准原因: "该日期位于 post-info 区域，紧邻标题，符合博客文章发布时间的典型位置特征，且早于其他候选日期，排除正文事件或活动日期。"
发布时间校准置信度: "1"
发布时间候选数量: 20
发布时间严格候选数量: 6
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T12:35:28+08:00"
发现时间: "2026-07-20T11:40:00+08:00"
入库时间: "2026-07-20T04:38:20.889Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://developer.nvidia.com/blog/"
匹配关键词:
  - "GPU"
  - "Vera Rubin"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 68
AI分档: "召回候选"
AI质检状态: "不通过"
AI打分理由: "NVIDIA官方发布，涉及AI工厂软件栈与Vera Rubin NVL72硬件协同，但核心聚焦软件平台（DSX OS）而非超节点硬件架构/供电/散热/互连等物理层细节，技术展开偏软件调度与运维，硬件参数缺失，商业落地信息有限。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T12:38:20+08:00"
AI主题相关性: 14
AI来源权威性: 15
AI新颖性: 16
AI技术细节: 10
AI商业部署信号: 8
AI完整性: 5
图片摘要:
  - "✓ ./assets/img-d4b26960.jpg | infographic | 展示AI工厂物理机架与DSX OS软件界面的结合，体现软硬件协同管理概念。"
  - "★ ./assets/img-57bcf34d.webp | diagram | NVIDIA DSX AI Factory平台架构图，展示从DSX OS软件层到参考硬件设计的全栈结构。"
  - "★ ./assets/img-d9e7228a.webp | diagram | DSX平台组件交互图，展示DSX Flex/MaxLPS与电网、BMS及Vera Rubin硬件的协同工作流。"
  - "✓ ./assets/img-42b7f730.webp | screenshot | Fleet Intelligence Dashboard界面，展示GPU/CPU利用率、内存状态等集群监控数据。"
  - "✗ ./assets/img-bab70b95.webp | ad | 会议注册广告，与正文技术内容无关"
  - "✗ ./assets/img-9d3a9443.webp | ad | 会议宣传广告，与正文技术内容无关"
采集批次: "2026年7月20日11点39分58秒"
采集批次ID: "20260720-113958-114"
去重键: "https://developer.nvidia.com/blog/nvidia-dsx-os-delivers-open-modular-software-for-operating-ai-factories-at-scale"
---

AI is now essential infrastructure, powered by AI factories that generate intelligence in the form of tokens. As demand grows, these factories must scale faster, operate more efficiently, and lower the cost of intelligence across the [five-layer stack](https://blogs.nvidia.com/blog/ai-5-layer-cake/): energy, chips, infrastructure, models, and applications.

[NVIDIA DSX](https://www.nvidia.com/en-us/data-center/products/dsx/) platform provides the complete playbook for designing, simulating, building, and operating AI factories, aligning every layer of the stack across compute, software, facilities, and partner technologies through a common co-designed architecture.

The DSX platform now includes [DSX OS software](https://docs.nvidia.com/dsx/home#dsx-os) to accelerate AI factory deployments and improve operational efficiency. DSX OS includes open source, modular software components and related NVIDIA technologies purpose-built for operating and scaling multi-tenant AI factories.

Together, DSX OS components enable NVIDIA DSX’s AI factory ecosystem to adopt the latest in agentic AI infrastructure software across the full stack, improving [tokens per watt](https://blogs.nvidia.com/blog/revenue-potential-ai-factories/) and lowering token cost, accelerating deployment, and strengthening operational reliability and resiliency.

![Architecture diagram showing NVIDIA DSX OS within the larger NVIDIA DSX platform across hardware, facilities, software, simulation, resiliency, and security layers](./assets/img-57bcf34d.webp)

Figure 1: NVIDIA DSX OS software in the DSX platform. DSX OS provides the open-source software for AI factory operations

## Why DSX OS matters to the AI factory ecosystem

AI factories must perform optimally in order to maximize the number of tokens they produce relative to the watts they consume, and bring real value to the operators.

In order to achieve this, [the complex network of components](https://developer.nvidia.com/blog/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt/) that goes into operating AI workloads at scale across datacenters must function in close harmony, requiring coordination across chips; systems; facilities infrastructure such as building management controls, cooling, and power distribution units; the power grid; the software and partner technologies running all of these; and the AI platforms and services running on top.

DSX OS software is designed for this entire ecosystem of components and provides a comprehensive set of open and extensible technologies and capabilities that can be integrated and adopted into existing platforms and software.

These capabilities have been designed and optimized around a common architecture, enabling all of the components involved to work together to deliver on three main outcomes that drive AI factory economics:

### 1) Faster time to revenue

NVIDIA builds and operates infrastructure and platform software on [NVIDIA DGX Cloud](https://www.nvidia.com/en-us/data-center/dgx-cloud), and now this software is being released as open source. NVIDIA ecosystem partners can leverage these components to deliver AI services rather than rebuild from scratch, eliminating months of custom development.

### 2) Better efficiency

Power is the limiting factor in an AI factory, and DSX connects power and grid behavior as part of the platform rather than as a facilities concern separated from the rest of the AI infrastructure. With DSX software, AI factories can run up to 40% more GPUs at peak energy efficiency within a fixed power budget, with minimal impact on inference workload performance.

### 3) Higher reliability and resiliency

AI factories run continuous large-scale workloads through hardware faults, grid events, and operational changes. DSX OS shifts cluster operations from reactive alerting to automated remediation, keeps runtime versions consistent across regions, and gives operators fleet-wide visibility.

## How DSX OS enables gigawatt-scale AI factories

The open source, modular components in DSX OS provide the foundational technologies for building and operating AI factories, and are designed to solve challenges unique to operating AI workloads efficiently and reliably at gigawatt scale.

They do so by providing a co-designed set of core capabilities, including (but not limited to) standardized communication, power and efficiency optimization, provisioning and lifecycle operations, health monitoring and remediation, and intelligent platform services.

More details about how DSX OS provides these capabilities follows:

### Standardized communication across the data center, enabled for agentic interfaces

An AI factory spans compute, networking, power, and cooling systems that all need to interoperate seamlessly. [DSX Exchange](http://github.com/NVIDIA/dsx-exchange) bridges these components with an MQTT-based IT/OT communication hub that makes facility-level signals such as grid events, thermal data, and power anomalies, visible to the software managing the rest of the AI factory, enabling components such as DSX Flex, MaxLPS, and partner software to react to each other’s state in real time, improving coordination and efficiency

DSX OS software components across the full DSX stack will also provide MCP servers for provisioning, networking, observability, and more. Using these MCP servers, AI agents can discover the entire operational surface of the factory as a unified tool catalog, enabling them to interface across every system and perform cross-domain correlation. With an agentic AI factory, operators can easily connect a GPU health event with a thermal anomaly, or a network issue to a performance issue, or other potential scenarios.

![A simplified diagram showing the connections between DSX Exchange, DSX Flex, DSX MaxLPS, provisioning systems such as NVIDIA Infra Controller, facilities components such as Building Management Systems, the power grid, third-party and partner software such as Emerald AI and Phaidra, and the Vera Rubin NVL 72 hardware](./assets/img-d9e7228a.webp)

Figure 2. DSX Exchange coordinates communication within the AI factory, including grid signals from DSX Flex, facilities-level signals, power policies to and from DSX MaxLPS, provisioning systems like NVIDIA Infra Controller, and more

### Power and efficiency optimization

Static power allocation strands capacity, reactive cooling creates thermal oscillations, and disconnected IT/OT systems make grid events a manual fire drill. DSX MaxLPS includes software that treats power as a programmable resource by dynamically enforcing policies at the GPU, rack, cooling, and workload level, enabling AI factories to recover stranded power to run additional compute at optimal utilization. DSX Flex extends this beyond the factory walls, with libraries for connecting workloads to grid services so AI factories can automatically adapt to demand response, load shedding, and renewable energy availability.  
  
Partners including CoreWeave, Firmus, Lambda, Nscale, and [Phaidra](https://www.phaidra.ai/) are deploying MaxLPS, while [Emerald AI](https://www.emeraldai.co/), ENGIE, Silicon Valley Power, and [UK National Grid](https://www.ngpartners.com/stories/emerald-ai-whitepaper) are leveraging DSX Flex.

### Provisioning and multi-tenant lifecycle operations

At scale, provisioning is a continuous workflow: nodes cycle through tenant assignments, hardware is replaced, and every transition must be auditable and secure. [NVIDIA Infra Controller (NICo)](https://docs.nvidia.com/infra-controller/documentation/home) makes this programmable with API-driven bare-metal lifecycle management and hardware-enforced tenant isolation through [NVIDIA BlueField DPUs](https://www.nvidia.com/en-us/networking/products/data-processing-unit/) and the [NVIDIA DOCA Platform Framework](https://www.nvidia.com/en-us/networking/products/software/doca/). [NVIDIA AI Cluster Runtime (AICR)](https://developer.nvidia.com/blog/validate-kubernetes-for-gpu-infrastructure-with-layered-reproducible-recipes/) complements this by capturing validated runtime configurations as version-locked recipes, eliminating the configuration drift that causes silent failures across large fleets.

IREN, OpenNebula Systems, Mirantis, Rafay, Red Hat, and Supermicro are among the partners integrating these components.

### Health monitoring and automation tooling

In a large GPU fleet, hardware degradation is a daily occurrence, and the traditional alert-page-investigate cycle is too manual for minimizing impact on workloads. [NVIDIA NVSentinel](https://developer.nvidia.com/blog/automate-kubernetes-ai-cluster-health-with-nvsentinel/) provides Kubernetes-native GPU fault detection and automated remediation, cordoning unhealthy compute nodes and draining workloads in seconds rather than minutes or hours. [NVIDIA Fleet Intelligence](https://developer.nvidia.com/blog/introducing-nvidia-fleet-intelligence-for-real-time-gpu-fleet-visibility-and-optimization/) provides fleet-wide visibility, integrity verification, and health monitoring across global deployments.  
  
Lambda is an early adopter of Fleet Intelligence.

![Screenshot of the Fleet Intelligence dashboard that summarizes fleet wide aggregations of data such as GPU and memory utilization as well as total GPUs in an up state](./assets/img-42b7f730.webp)

Figure 3. The NVIDIA Fleet Intelligence dashboard summarizes fleet-wide aggregations of data such as GPU and memory utilization as well as total GPUs in an up state

### Intelligent AI workload scheduling and platform services

AI workloads need more than GPU access; they need topology-aware intelligent scheduling, distributed inference, and production APIs. [KAI Scheduler](https://developer.nvidia.com/blog/nvidia-open-sources-runai-scheduler-to-foster-community-collaboration/) and [NVIDIA Run:ai](https://www.nvidia.com/en-us/software/run-ai) provide GPU-aware workload placement with fractional allocation and hierarchical quotas. [NVIDIA Dynamo](https://developer.nvidia.com/dynamo) and [NVIDIA Grove](https://developer.nvidia.com/grove) deliver distributed inference serving with disaggregated prefill/decode and per-stage autoscaling. [NVIDIA Cloud Functions (NVCF)](https://developer.nvidia.com/dgx-cloud/nvcf) ties it together with unified APIs across inference, fine-tuning, and batch workloads with built-in multi-tenancy.  
  
Partners including Aible, Beyond AI, Bhashini, Crusoe, DCAI, Mirantis, Nebius, Rafay, Sarvam, Simplismart, Spectro Cloud, vCluster, Vultr, and Yotta are using many of these components in production.

## Getting started

DSX OS components are available on GitHub and designed for incremental adoption and integration with existing software stacks.

Start with the component that addresses your most immediate requirements, and build from there, leveraging the capabilities and technologies provided to accelerate your AI factory deployment and improve operational efficiency.

Some examples are provided below:

- IT/OT communications: [DSX Exchange](https://github.com/NVIDIA/dsx-exchange)
- Bare-metal lifecycle management and tenant isolation: [NVIDIA Infra Controller](https://github.com/NVIDIA/infra-controller-core) and [DOCA Platform Framework](https://github.com/NVIDIA/doca-platform/)
- Fleet visibility, health, and integrity: [NVIDIA Fleet Intelligence](https://github.com/NVIDIA/fleet-intelligence-agent/)
- Unified AI inference APIs: [NVIDIA Cloud Functions](https://github.com/NVIDIA/nvcf)

Review [NVIDIA DSX documentation](https://docs.nvidia.com/dsx) for more details about all of the components of DSX OS, implementation and reference design guides, quickstarts, and integration guidance.

![图片](./assets/img-d4b26960.jpg)

![Architecture diagram showing NVIDIA DSX OS within the larger NVIDIA DSX platform across hardware, facilities, software, simulation, resiliency, and security layers](./assets/img-57bcf34d.webp)_**Figure 1: NVIDIA DSX OS software in the DSX platform.** DSX OS provides the open-source software for AI factory operations_

![A simplified diagram showing the connections between DSX Exchange, DSX Flex, DSX MaxLPS, provisioning systems such as NVIDIA Infra Controller, facilities components such as Building Management Systems, the power grid, third-party and partner software such as Emerald AI and Phaidra, and the Vera Rubin NVL 72 hardware](./assets/img-d9e7228a.webp)_Figure 2. DSX Exchange coordinates communication within the AI factory, including grid signals from DSX Flex, facilities-level signals, power policies to and from DSX MaxLPS, provisioning systems like NVIDIA Infra Controller, and more_

![Screenshot of the Fleet Intelligence dashboard that summarizes fleet wide aggregations of data such as GPU and memory utilization as well as total GPUs in an up state](./assets/img-42b7f730.webp)_Figure 3. The NVIDIA Fleet Intelligence dashboard summarizes fleet-wide aggregations of data such as GPU and memory utilization as well as total GPUs in an up state_

- ![图片](./assets/img-bab70b95.webp)

- ![图片](./assets/img-9d3a9443.webp)
