---
格式版本: 2
标题: "Cosmos World Foundation Models Openly Available to Physical AI Developers | NVIDIA Blog"
原文链接: "https://blogs.nvidia.com/blog/cosmos-world-foundation-models/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期来自HTML中的time标签，class包含nvidia-article-date，且位于标题附近，符合文章发布时间的特征。"
发布时间校准置信度: "1"
发布时间候选数量: 32
发布时间严格候选数量: 8
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T11:44:22+08:00"
发现时间: "2026-07-20T09:25:08+08:00"
入库时间: "2026-07-20T03:48:31.375Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=C-Link"
匹配关键词:
  - "C-Link"
  - "GPU"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 35
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "文章主要介绍NVIDIA Cosmos世界基础模型平台，用于物理AI（机器人、自动驾驶）开发，属于AI软件/模型层面，未涉及超节点、AI Rack、机柜级系统、供电、散热、高速互连等硬件基础设施内容，与项目关注范围不符。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T11:48:31+08:00"
AI主题相关性: 5
AI来源权威性: 15
AI新颖性: 10
AI技术细节: 5
AI商业部署信号: 0
AI完整性: 0
图片摘要:
  - "✗ ./assets/img-83d2f103.jpg | diagram | 图片主题为AI基础设施能效（Performance per Watt），与正文Cosmos世界基础模型及物理AI开发主题无关。"
  - "✗ ./assets/img-1f6575f1.png | diagram | 图片主题为NVIDIA Vera CPU架构及单线程性能，与正文Cosmos世界基础模型主题无关。"
  - "✗ ./assets/img-3511966c.jpg | photo | 图片为NVIDIA公司标志及建筑实拍，属于品牌宣传或通用配图，未展示Cosmos模型或物理AI相关技术细节。"
  - "✗ ./assets/img-e47e99dc.png | infographic | 图片主题为推理软件栈与Token成本，与正文Cosmos世界基础模型发布及物理AI开发主题无关。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/cosmos-world-foundation-models"
---

*Editor’s note: This post was updated on Friday, Jan. 10, with Best of CES Awards results.*

[NVIDIA Cosmos](https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-world-foundation-model-platform-to-accelerate-physical-ai-development), a platform for accelerating [physical AI](https://www.nvidia.com/en-us/glossary/physical-ai/) development, introduces a family of [world foundation models](https://www.nvidia.com/en-us/glossary/world-models/) — neural networks that can predict and generate physics-aware videos of the future state of a virtual environment — to help developers build next-generation robots and autonomous vehicles (AVs).

World foundation models, or WFMs, are as fundamental as large language models. They use input data, including text, image, video and movement, to generate and simulate virtual worlds in a way that accurately models the spatial relationships of objects in the scene and their physical interactions.

[Announced at CES](https://www.ces.tech/schedule/nvidia-keynote/), NVIDIA is making available the first wave of Cosmos WFMs for physics-based simulation and synthetic data generation — plus state-of-the-art tokenizers, guardrails, an accelerated data processing and curation pipeline, and a framework for model customization and optimization.

Cosmos won Best AI and Best Overall accolades from the [Best of CES Awards](https://www.cnet.com/tech/these-are-the-official-2025-best-of-ces-winners-awarded-by-cnet-group/) by the CNET Group, awards partner for the Consumer Technology Association, which produces CES.

Researchers and developers, regardless of their company size, can freely use the Cosmos models under NVIDIA’s permissive open model license that allows commercial usage. Enterprises building AI agents can also use new open [NVIDIA Llama Nemotron and Cosmos Nemotron models](https://blogs.nvidia.com/blog/nemotron-model-families), unveiled at CES.

The openness of Cosmos’ state-of-the-art models unblocks [physical AI](https://www.nvidia.com/en-us/glossary/physical-ai/) developers building robotics and AV technology and enables enterprises of all sizes to more quickly bring their physical AI applications to market. Developers can use Cosmos models directly to generate physics-based synthetic data, or they can harness the [NVIDIA NeMo framework](https://github.com/NVIDIA/NeMo/tree/main/nemo/collections/diffusion) to fine-tune the models with their own videos for specific physical AI setups.

Physical AI leaders — including robotics companies 1X, Agility Robotics and XPENG, and AV developers Uber and Waabi — are already working with Cosmos to accelerate and enhance model development.

Developers can preview the first Cosmos [autoregressive](https://build.nvidia.com/nvidia/cosmos-1_0-autoregressive-5b) and [diffusion](https://build.nvidia.com/nvidia/cosmos-1_0-diffusion-7b) models on the [NVIDIA API catalog](https://build.nvidia.com/explore/discover), and download the family of models and fine-tuning framework from the [NVIDIA NGC catalog](https://catalog.ngc.nvidia.com/) and [Hugging Face](https://huggingface.co/collections/nvidia/cosmos-world-models-6751e884dc10e013a0a0d8e6).

NVIDIA Cosmos: A World Foundation Model Platform for Physical AI - YouTube

NVIDIA 2.23M subscribers

## World Foundational Models for Physical AI

Cosmos world foundation models are a suite of open diffusion and autoregressive transformer models for physics-aware video generation. The models have been trained on 9,000 trillion tokens from 20 million hours of real-world human interactions, environment, industrial, robotics and driving data.

The models come in three categories: Nano, for models optimized for real-time, [low-latency inference](https://developer.nvidia.com/blog/tag/low-latency-inference/) and edge deployment; Super, for highly performant baseline models; and Ultra, for maximum quality and fidelity, best used for distilling custom models.

When paired with [NVIDIA Omniverse](https://www.nvidia.com/en-us/omniverse/) 3D outputs, the diffusion models generate controllable, high-quality synthetic video data to bootstrap training of robotic and AV perception models. The autoregressive models predict what should come next in a sequence of video frames based on input frames and text. This enables real-time next-token prediction, giving physical AI models the foresight to predict their next best action.

Developers can use Cosmos’ open models for text-to-world and video-to-world generation. Versions of the diffusion and autoregressive models, with between 4 and 14 billion parameters each, are available now on the NGC catalog and [Hugging Face](https://huggingface.co/collections/nvidia/cosmos-world-models-6751e884dc10e013a0a0d8e6).

Also available are a 12-billion-parameter upsampling model for refining text prompts, a 7-billion-parameter video decoder optimized for augmented reality, and guardrail models to ensure responsible, safe use.

To demonstrate opportunities for customization, NVIDIA is also releasing fine-tuned model samples for vertical applications, such as generating multisensor views for AVs.

## Advancing Robotics, Autonomous Vehicle Applications

Cosmos world foundation models can enable [synthetic data generation](https://www.nvidia.com/en-us/use-cases/synthetic-data/) to augment training datasets, simulation to test and debug physical AI models before they’re deployed in the real world, and reinforcement learning in virtual environments to accelerate [AI agent learning](https://blogs.nvidia.com/blog/what-is-agentic-ai/).

Developers can generate massive amounts of controllable, physics-based synthetic data by conditioning Cosmos with composed 3D scenes from NVIDIA Omniverse.

Waabi, a company pioneering generative AI for the physical world, starting with autonomous vehicles, is evaluating the use of Cosmos for the search and curation of data for AV software development and simulation. This will further accelerate the company’s industry-leading approach to safety, which is based on Waabi World, a generative AI simulator that can create any situation a vehicle might encounter with the same level of realism as if it happened in the real world.

In robotics, WFMs can generate synthetic virtual environments or worlds to provide a less expensive, more efficient and controlled space for robot learning. Embodied AI startup Hillbot is boosting its data pipeline by using Cosmos to generate terabytes of high-fidelity 3D environments. This AI-generated data will help the company refine its robotic training and operations, enabling faster, more efficient robotic skilling and improved performance for industrial and domestic tasks.

In both industries, developers can use NVIDIA Omniverse and Cosmos as a multiverse simulation engine, allowing a physical AI policy model to simulate every possible future path it could take to execute a particular task — which in turn helps the model select the best of these paths.

Data curation and the training of Cosmos models relied on thousands of NVIDIA GPUs through [NVIDIA DGX Cloud](https://www.nvidia.com/en-us/data-center/dgx-cloud/), a high-performance, fully managed AI platform that provides accelerated computing clusters in every leading cloud.

Developers adopting Cosmos can use DGX Cloud for an easy way to deploy Cosmos models, with further support available through the [NVIDIA AI Enterprise](https://www.nvidia.com/en-us/data-center/products/ai-enterprise/) software platform.

## Customize and Deploy With NVIDIA Cosmos

In addition to foundation models, the [Cosmos platform](http://www.nvidia.com/en-us/ai/cosmos) includes a data processing and curation pipeline powered by [NVIDIA NeMo Curator](https://developer.nvidia.com/nemo-curator) and optimized for NVIDIA data center GPUs.

Robotics and AV developers collect millions or billions of hours of real-world recorded video, resulting in petabytes of data. Cosmos enables developers to process 20 million hours of data in just 40 days on [NVIDIA Hopper GPUs](https://www.nvidia.com/en-us/data-center/technologies/hopper-architecture/), or as little as 14 days on [NVIDIA Blackwell GPUs](https://www.nvidia.com/en-us/data-center/technologies/blackwell-architecture/). Using unoptimized pipelines running on a CPU system with equivalent power consumption, processing the same amount of data would take over three years.

The platform also features a suite of powerful video and image tokenizers that can convert videos into tokens at different video compression ratios for training various [transformer models](https://blogs.nvidia.com/blog/what-is-a-transformer-model/).

The Cosmos tokenizers deliver 8x more total compression than state-of-the-art methods and 12x faster processing speed, which offers superior quality and reduced computational costs in both training and [inference](https://www.nvidia.com/en-us/solutions/ai/inference/). Developers can access these tokenizers, available under NVIDIA’s open model license, via [Hugging Face](https://huggingface.co/nvidia/Cosmos-Tokenizer-CI8x8) and [GitHub](https://github.com/NVIDIA/Cosmos-Tokenizer).

Developers using Cosmos can also harness model training and fine-tuning capabilities offered by [NeMo framework](https://docs.nvidia.com/nemo-framework/user-guide/latest/overview.html), a GPU-accelerated framework that enables high-throughput AI training.

## Developing Safe, Responsible AI Models

Now available to developers under the NVIDIA Open Model License Agreement, Cosmos was developed in line with NVIDIA’s [trustworthy AI](https://www.nvidia.com/en-us/ai-data-science/trustworthy-ai/) principles, which include nondiscrimination, privacy, safety, security and transparency.

The Cosmos platform includes Cosmos Guardrails, a dedicated suite of models that, among other capabilities, mitigates harmful text and image inputs during preprocessing and screens generated videos during postprocessing for safety. Developers can further enhance these guardrails for their custom applications.

Cosmos models on the [NVIDIA API catalog](https://build.nvidia.com/explore/discover) also feature an inbuilt watermarking system that enables identification of AI-generated sequences.

NVIDIA Cosmos was developed by [NVIDIA Research](https://www.nvidia.com/en-us/research/). Read the research paper, “ [Cosmos World Foundation Model Platform for Physical AI](https://arxiv.org/abs/2501.03575),” for more details on model development and benchmarks. Model cards providing additional information are available on [Hugging Face](https://huggingface.co/collections/nvidia/cosmos-world-models-6751e884dc10e013a0a0d8e6).

*Learn more about world foundation models in an* [*AI Podcast episode*](https://blogs.nvidia.com/blog/world-foundation-models-advance-physical-ai/) *that features Ming-Yu Liu, vice president of research at NVIDIA.*

[*Get started*](https://developer.nvidia.com/cosmos) *with NVIDIA* *Cosmos.*

*Visit our* *[Cosmos Cookbook](https://nvda.ws/4qevli8)* *for step-by-step workflows, technical recipes, and concrete examples for building, adapting, and deploying Cosmos WFMs, or* *[join our community](https://discord.gg/u23rXTHSC9)* *to learn with peers.*

*See* [*notice*](https://nam11.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.nvidia.com%2Fen-eu%2Fabout-nvidia%2Fterms-of-service%2F&data=05%7C02%7Clpham%40nvidia.com%7C50be0315fe044aff6e4208dd246c5d29%7C43083d15727340c1b7db39efd9ccc17a%7C0%7C0%7C638706770021365260%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=NMjf0ustmoQOEXvAwyKwtNLu8m%2BXF%2FYzs2BUq2HQBgY%3D&reserved=0) *regarding software product information.*
