---
格式版本: 2
标题: "How to Build End‑to‑End Physical AI Systems for Humanoid Robots S81478 | GTC San Jose 2026"
原文链接: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/"
发布日期: "2026-07-30"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:scrape:provider_published_at"
发布时间证据: "provider publishedAt: 2026-07-30"
发布时间校准原因: "规则确认唯一严格发布时间，来源 scrape:provider_published_at"
发布时间校准置信度: "high"
发布时间候选数量: 6
发布时间严格候选数量: 2
发布时间原页读取状态: "原页面已读取"
发布时间未找到原因: ""
发布时间校准时间: "2026-07-31T21:46:37+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-07-31T21:46:24+08:00"
入库时间: "2026-07-31T13:46:37.416Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.nvidia.com/gtc/"
匹配关键词:
  - "GPU"
  - "roadmap"
  - "deployment"
  - "performance"
  - "latency"
相关厂家:
  - "NVIDIA"
  - "Microsoft"
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 18
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为机器人物理AI系统构建，与超节点/AI Rack/机柜级AI基础设施无关，不涉及GPU、互连、供电、散热或量产部署等核心主题，属于低价值无关内容。"
AI质检模型: "deepseek-v4-flash"
AI质检时间: "2026-07-31T21:47:55+08:00"
AI主题相关性: 2
AI来源权威性: 12
AI新颖性: 2
AI技术细节: 1
AI商业部署信号: 0
AI完整性: 1
采集批次: "2026年7月31日21点46分24秒"
采集批次ID: "20260731-214624-519"
去重键: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478"
---

Title: How to Build End‑to‑End Physical AI Systems for Humanoid Robots S81478 | GTC San Jose 2026

URL Source: https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/

Published Time: Thu, 30 Jul 2026 19:40:37 GMT

Markdown Content:
Visit your regional NVIDIA website for local content, pricing, and where to buy partners specific to your country.

[Continue](https://www.nvidia.com/)

[**GTC Berlin** October 20–22](https://www.nvidia.com/en-eu/gtc/) | [**GTC 2027** March 15–18](https://www.nvidia.com/gtc/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#) 
*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)      
*   [](https://www.nvidia.com/en-us/account/)
*   [Log In](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)[LogOut](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

PLATFORMS

other links

[](https://www.nvidia.com/gtc/)

 Keynote 
*   [Keynote](https://www.nvidia.com/gtc/keynote/)
*   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

 Explore 
*   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
*   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
*   [Speakers](https://www.nvidia.com/gtc/speakers/)
*   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
*   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

[Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)

 More 
*   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
*   [Contact Us](https://www.nvidia.com/gtc/contact/)
*   [FAQ](https://www.nvidia.com/gtc/faq/)
*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*    Keynote 
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*    Explore 
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*    More 
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/# "Menu")

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/# "Menu")

*   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81478/#)
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

[Video 1](blob:https://www.nvidia.com/afda1fb8-524f-4e56-b43b-2f106ca593fd)

Play

Play

Play on TV

00:00

Play

Seek 10 seconds backwards

Seek 10 seconds forward

00:00 / 01:17:13

Mute

Click to volume control Use the arrows to control the volume

Enable Captions

Settings

Play on TV

Turn on Picture in picture

Show Full screen

Transcript Powered by AI

X

X

00:12

Okay, perfect, yeah.

00:13

So we are entering the era of physical AI.

00:16

Now the idea of intelligent machines, robots, autonomous

00:20

vehicles, and smart factories isn't new.

00:23

We have been building them for decades.

00:26

But what's fundamentally changing is how we build them.

00:29

Traditionally, each task required its own dedicated model, one

00:34

for perception, then one for planning, and one for control.

00:38

So each trained separately and then stitched together by hand.

00:43

That made scaling incredibly difficult.

00:45

Today, what we are witnessing is that a shift towards a unified

00:49

end-to-end foundation model.

00:52

Models that can directly map raw sensor data to real-world actions.

00:57

The benefits are significant.

00:59

These models generalize far better to unseen situations,

01:03

respond faster in real-time, and are much easier to adapt across

01:07

different robots and environments.

01:11

So for a long time, robotics has relied on these specialist

01:14

models, the kind that you see in factories and warehouses right now.

01:19

So these models had to be programmed by hand for rigid,

01:23

controlled environments, but they were very proficient

01:27

at executing specific tasks with speed and precision.

01:31

Now, we are seeing the rise of generalist models, models

01:35

that learn tasks rather than being explicitly programmed.

01:39

They have broad scope, but their proficiency on any single

01:42

task is still limited.

01:44

Our goal with physical AI is to bridge that gap, to

01:47

scale intelligence to a point where we can develop generalist

01:51

specialist models that combine both broad scope and deep proficiency.

02:02

But physical AI is hard to develop.

02:04

First, real-world data is inherently multimodal.

02:08

You're dealing with cameras, sensors, LIDAR, tactile feedback,

02:13

and force feedback.

02:15

Second, the data is extremely difficult to collect.

02:18

For robotics, you typically need teleoperation, which

02:21

requires expensive hardware, carefully controlled environments,

02:25

and constant human supervision.

02:28

Third, deployment at edge is a real problem, a hard problem in itself.

02:34

You're working under very tight latency constraints

02:38

with limited compute budgets.

02:40

And then once you train it, testing is equally painful.

02:44

It's slow, it's risky, and it's hard to reproduce at scale.

02:50

So what you're seeing here are real videos from our lab

02:53

where you can see the robot crashing and falling down

02:57

while we are testing new policies.

02:58

So these tests can be quite dangerous and very expensive.

03:08

Yeah, so today we are going to talk about the different

03:10

workflows, technologies, and some exciting research that

03:14

we have been doing at NVIDIA.

03:16

We have divided up our presentation into three parts.

03:19

Workflows for data generation, training, and deployment.

03:22

Let's start with data generation.

03:27

So we have had these different eras in AI over the years.

03:32

The big unlock for LLMs or generative AI has been the

03:36

fact that they have been trained with trillions of tokens from

03:39

the whole internet.

03:41

And with AGENTIC AI and reasoning, we allowed AI to self-reflect on

03:45

its answer and improve them using a process called test time scaling,

03:48

using reinforcement learning.

03:51

All these AI's are a result of utilizing large amounts of

03:56

compute over large amounts of data.

04:00

Now in robotics, we don't have any such data to train models.

04:04

At best, what we have is a few hundred thousand hours

04:07

of tele-op data, and most of it is very task-specific.

04:12

So now the question becomes, how do we convert a problem that

04:18

is bounded by data into a problem that is bounded by computer?

04:26

We believe we can solve this using two foundational technologies.

04:30

The first is simulation, and second is world foundation models.

04:35

NVIDIA Omniverse is our simulation platform with iSEC Sim and iSEC

04:39

Lab, our simulators for robotics.

04:42

NVIDIA Cosmos is our world foundation model platform.

04:46

With both these platforms, we can generate large, diverse

04:50

data sets which can be used to train physical AI or even give

04:54

AI the ability to experience the data encoded in these world models.

04:59

And these are only limited by the amount of compute we have.

05:05

So at NVIDIA, we call this the data pyramid for robots.

05:09

At the top, we have real-world data, small, expensive, and active.

05:13

A robot might collect like 24 hours of it per day.

05:18

In the middle, we have the synthetic data, which is

05:21

infinite in principle.

05:22

We can collect like GBs per day.

05:25

GB is per GPU per day, and at the base, we have the web

05:28

data, which is unstructured and multimodal and exabytes.

05:32

The goal is to grow the middle layer until it becomes a dominant

05:36

source of training.

05:37

When synthetic data surpasses web-scale data, that's when

05:40

robots can truly learn to become generalized for every task.

05:44

That's our vision for physical AI.

05:49

Okay, I want to first talk about Cosmos, our world

05:51

foundation model platform.

05:53

So there are four components.

05:55

First is the foundation models, predict, transfer, and reason.

05:59

Predict and transfer are world generation models,

06:01

while reason is a VLM with chain of thought reasoning.

06:05

The second is the frameworks that we have.

06:07

Cosmos Curator is our data curation and processing pipeline for

06:11

processing large amounts of data.

06:14

Cosmos RL enables you to do large-scale RL training, and

06:18

Cosmos Evaluator is our automated evaluation and grading system

06:22

for synthetic video output generated by the Cosmos models.

06:26

Third, we have the different blueprints.

06:28

These are reference implementations that demonstrate how to compose

06:32

Cosmos models together for specific use cases.

06:35

And then we have the post-training and inference scripts, and Cosmos

06:39

Cookbook, which is our guide for working with the Cosmos ecosystem.

06:56

Okay, so let's start with PREDICT, which is our base

07:01

world foundation model.

07:02

PREDICT is designed for simulating and predicting future states

07:08

of the world as a video.

07:10

It takes text, image, or video as input and generates the

07:13

future world state as a video.

07:15

In the latest version of Predict, Predict 2.5, we can generate up

07:18

to 30 seconds of the future state.

07:21

This model is open source, weights are on Hugging Face,

07:23

with post-training scripts available on GitHub.

07:26

So the goal with Predict is to use it as a world foundation model,

07:30

which you can then post-train to create specialized AI models.

07:34

So let's actually go over one of this here.

07:37

So DreamDojo is a research experiment we did using

07:41

the Predict model.

07:42

So the way it works is that it's a foundation model that

07:45

was built on top of PREDICT.

07:47

It was pre-trained on 44,000 hours of egocentric human video using

07:52

latent actions as proxy labels.

07:55

So the model learns the dynamics of the physical

07:58

interaction without needing robot-specific data upfront.

08:03

After this pre-training, we then post-train the model

08:07

using a small teleoperated dataset.

08:10

This step teaches the model to target the robot's action space,

08:15

and this is what enables it to generate realistic robot rollouts

08:19

with minimal real-world data.

08:21

Now the base model of Predict requires like 35 denoising

08:24

steps, so it's very slow for real-time inference.

08:28

So we do a distillation where we distillate down to four steps,

08:32

achieving real-time inference at almost 10 frames per second.

08:36

So now let me show you the results from an experiment we did, where

08:39

we took the DreamDojo model and we post-trained with a small data

08:43

set of G1, unitary G1 teleop data.

08:47

So, after post-training, you can see that the model can

08:50

simulate the robot interacting with objects that it has never seen

08:54

during the training, like a paper or a flower book while maintaining

09:01

the physical possible behavior.

09:05

Beyond new objects, the model is also able to generalize to entirely

09:09

unseen environments that was not present in the teleop data set.

09:14

You can also do real-time teleoperation, where the model

09:18

is generated on the fly based on action conditioning it

09:23

got from the teleop device.

09:27

So next, let's look at another workflow called Cosmos Policy.

09:32

So Cosmos can generate diverse robot videos given

09:36

the input images and text.

09:38

Now the question is, can we post in these models to also generate

09:42

actions that control a robot?

09:45

So Cosmos policy, what it does is that it turns predict into

09:49

a fully-functional robot policy with zero architectural changes.

09:53

So the policy simply encodes the actions, the robot state, and

09:59

the value function as additional latent frames and injects them

10:02

into the same sequence that the model already knows how to process.

10:06

So the model doesn't distinguish between predicting

10:10

the next video frame and predicting the next action.

10:14

From the perspective of the model, it's just predicting

10:17

the next latent frame.

10:18

And the architecture completely remains unchanged.

10:21

But now, it learns to control a robot.

10:24

It's really cool. Let me show you some of the results

10:26

we have with this.

10:29

So here we can see the Cosmos policy can perform language

10:33

condition pick and place following the user's prompt.

10:42

Here, Cosmos Policy is performing a long-horizon,

10:45

contact-rich manipulation task, where it's folding

10:47

like a t-shirt with many tips.

10:49

The cool part about this whole thing is that it was only trained

10:52

with 15 training demonstrations.

10:55

It's incredible to understand that because Predict wasn't

10:57

trained on large amounts of robot teleoperation data.

11:01

It didn't understand the concept of controlling a robot or

11:04

anything like that and yet with only 15 training demonstrations

11:08

the model was able to understand it and be able to control

11:10

it and be able to do this task.

11:15

Here, Cosmos' policy performs consecutive object placements,

11:20

handling the training data set with high-action

11:22

multimodality, given the random alterations to the left arm.

11:27

So this is a cool one.

11:29

So it's doing a high-precision and it's carefully high-precision

11:33

manipulation task, where you can see that it opens this block

11:35

back and puts the candy inside.

11:37

This simple task is actually quite complex, and it's amazing

11:40

that the model was able to do it.

11:43

So with all this,

11:45

You can understand that PREDICT is demonstrating that it's

11:47

a single world foundation model that can serve as a

11:50

powerful base for creating specialized physical AI models.

11:56

One that is flexible enough to learn from diverse data

11:59

and to create entirely new neural trajectories while

12:02

also being capable enough to directly control a robot.

12:08

Next, let's look at transfer, which is built on top of predict.

12:12

So unlike predict, it doesn't generate new states, but instead

12:18

allows you to create variations to existing data.

12:22

It uses the control net architecture, which allows

12:24

it to preserve the structure of the video while selectively changing

12:28

things like scene composition, object placement, et cetera.

12:31

The way it works is that you provide a video that you want

12:34

to create variations, and then you also provide control videos

12:38

like edge, seg, vis, and depth to precisely create these variations.

12:45

So when you're doing real-to-real augmentation from a real video,

12:49

we can use these different models and algorithms that

12:51

we have to create the different modalities, like groundingDyno

12:54

plus SAM2 for second rule, or like depth anything for depth control.

13:00

Okay, now let's see transfer in action with a simple example.

13:04

So this is a standard manipulation task where the robot arm is picking

13:08

up what I believe is a lettuce and placing it inside a wire mesh.

13:13

So let's try to use transfer to do an object change where

13:18

it will change the lettuce into a light bulb.

13:23

So whenever we are using transfer, we need to ask

13:27

yourself two questions.

13:28

What do I have to preserve, and what do I need to change?

13:32

So in this case, we need to preserve pretty much everything

13:35

apart from the letters, and maybe even the hand that grabs

13:38

the letters, because Lightbulb will be shaped a bit different, and

13:42

you might have to change the hand a bit in order to be able to grab it.

13:45

All right.

13:47

So to preserve things structurally, we use the edge control modality.

13:53

And to preserve the colors, we use this.

13:56

The higher the control rate, the more aggressive the model

13:59

preserves things.

14:06

So then we use the second rule to change things.

14:12

But since we only want to change the letters and the

14:14

hand, we need to use a binary mask, which is our second

14:19

video that you're seeing, where the white pixels tell Cosmos where

14:23

to apply the control modality.

14:26

The last video that you're seeing is what the model internally

14:29

sees from the mask and the control modality that we gave.

14:34

Now, when you generate the video, you will get this,

14:37

where the letters are turned into a light bulb, while everything

14:41

else remains unchanged.

14:47

All right, next, let's look at a workflow using Transfer

14:51

for synthetic data generation for policy training.

14:54

So here is another experiment we did in NVIDIA

14:57

to demonstrate this workflow.

14:58

The objective here is to train a policy that can do a standard

15:03

manipulation task.

15:04

So the task here is that you pick up a bowl with one hand,

15:07

apple with the other, place the apple in the bowl, and

15:09

then put the bowl back on the table under varied visual conditions.

15:13

We collected 100 teleop data from the real robot and then

15:18

we created three training variants.

15:20

One is the base policy, which only has these 100 real demos,

15:24

no augmentations.

15:25

The second is a baseline policy, which has the same 100 demos

15:28

plus standard image augmentations like noise and jitter, etc.

15:33

And the third is the Cosmos augmented policy, which has

15:36

these 100 demos plus 5x Cosmos transfer This was generated data.

15:42

All other parameters were kept identical.

15:46

So these are some of the data that Cosmos Transfer generated.

15:52

Again, more generations.

15:58

Then we did a generalization test where we evaluated on 10

16:02

unseen scene configurations where we had things like new bowels,

16:06

new fruits, different backgrounds, lighting changes, distracted

16:10

objects on the table, and the drawers in the back were open.

16:14

We changed the tablecloths.

16:16

And then when we tested, you can see that the Cosmos augmented

16:18

policy succeeds 80% of the time, while the base policy only succeeds

16:26

on the training scene, while the baseline only succeeds 16 times.

16:30

The table below kind of shows the different results.

16:34

So the key takeaway here is that strong generalization

16:37

capabilities, you can achieve them through Cosmos Transfer.

16:41

So then we deployed it on the real robot.

16:43

And you can see that the ones at the bottom, which use the

16:48

Cosmos transfer augmented policy, succeed in these tasks.

16:57

So now let's talk about Cosmos Reason.

17:01

So traditionally, VLM models can understand what's happening

17:05

in an image or video, but they struggle to reason through the

17:09

unfamiliar or complex scenarios.

17:12

Cosmos Reason addresses that.

17:15

It's a reasoning VLM that combines visual understanding

17:18

with physical reasoning.

17:20

This model is available on Hugging Face and InfraScripts on GitHub.

17:24

So Cosmos recent can be used for data annotation in the

17:30

Cosmos curator pipeline.

17:32

It can accurately create dense captions that can annotate

17:38

videos for training datasets.

17:43

It can also be used to automatically critique

17:46

training videos that are generated from other Cosmos models.

17:50

OK, now let's jump back to simulation.

17:53

So it is important that the worlds we are constructing for physical

17:56

AI must accurately reflect reality.

17:59

And with neural reconstruction, we can bring the real world

18:02

directly into simulation.

18:04

So Omniverse NeurIQ is a set of APIs and libraries to generate

18:09

interactive 3D simulation from the real world data.

18:12

Here we are using a library called 3DGrad.

18:15

To generate Gaussian Splats from the video captured

18:18

through a cell phone.

18:19

This is our cafeteria in our office in Zurich.

18:22

And once we have trained our Splats, you can export it

18:24

as a USD file, and we can bring it to Isaac Sim.

18:31

Once it is in Isaac Sim, now you can use the Splat

18:34

to train your robots.

18:35

This provides an accurate visual fidelity which otherwise was

18:39

very hard to achieve in simulation.

18:45

So next is Isaac Teleop, our unified framework for

18:49

both simulated and real robot teleoperation.

18:53

The core problem it solves is integration.

18:55

So today, setting up teleoperation means that it requires significant

19:00

effort, where you have to do the integration work for

19:04

the different devices working together, like headsets, foot

19:08

pedals, motion trackers, and so on.

19:10

And you typically end up building separate pipelines

19:13

for sim and real.

19:15

Isaac Teleop unifies all of this into a single framework.

19:20

With Isaac Teleop, we now have a single framework that

19:23

integrates these different headsets and control devices out of the box,

19:27

provides standard interfaces for common retargeters, and supports

19:31

both 2D and 3D camera output.

19:33

And it works seamlessly with both real robots and simulated

19:37

robots in Isaac Lab.

19:41

So next is SOMA.

19:43

In robotics and animations, there are many different

19:46

digital models that are used to represent the human

19:49

body in 3D, each built by different research labs and using their

19:55

own skeleton structures and data formats, making them completely

20:00

incompatible with each other.

20:02

SOMA was created by NVIDIA as a universal translator, so

20:06

no matter where your human motion data comes from, for example,

20:10

like a video of someone moving, or a motion that was generated

20:15

from a test description, or an existing motion capture dataset,

20:19

like the The bone seed data set.

20:21

SOMA converts it all into a single shared format, which

20:24

can then be retargeted using the SOMA retargeter, so that the robot

20:30

can learn to replicate this motion.

20:31

For example, in this video that you're seeing, we use

20:34

the Kimodo model to generate, using With a text description

20:40

to generate a robot dancing motion.

20:43

Which then flows through Soma into the retargeter and maps

20:47

to the robot joints to get a reference trajectory.

20:50

Once you have a reference trajectory, now you can train

20:52

your policy in RL.

20:54

We used another NVIDIA technology called, I believe

20:58

it's called protomotions.

21:02

We also released a massive open source data set for

21:05

physical AI development.

21:07

So training these robots require large amounts of high quality data.

21:11

And we are open sourcing our own training data that we

21:14

used to help developers and researchers accelerate their work.

21:17

This is commercial-grade, pre-validated data that can

21:20

be used for model training and testing and validation.

21:24

The initial release has around 500,000 real and synthetic

21:28

trajectories for robotics, almost 2,000 hours of real driving data,

21:33

and 1,000 USD and SIM-ready assets.

21:36

We'll be continuously expanding this data set

21:38

as our development progresses.

21:41

OK, I know that was a lot of things, but we are at the

21:45

end of the first session.

21:46

For now to talk about training and deployment, I invite Edith.

22:00

Awesome. Hello, everyone.

22:01

Thanks, Akul, for that introduction.

22:03

My name is Edith, and I will be continuing on for the rest

22:07

of this presentation.

22:09

So Akul just gave an overview of all of the technology that we

22:13

have available for data generation.

22:16

Now, in order for my robot to do a particular task, what's next?

22:21

What follows?

22:22

Well, let's take a look at training.

22:24

We need to train a policy in order for our robot to

22:27

do a particular task that we want.

22:30

And let's see the tools that NVIDIA has to support this.

22:37

So there are three very common robot learning paradigms.

22:42

Imitation learning, where you're learning based on trying to mimic

22:47

demonstrations from an expert.

22:49

Reinforcement learning, where you're learning your policies

22:52

through trial and error.

22:53

With reward signals, and there's no human demonstration kind

22:57

of required here.

22:58

And then finally, vision language action models that take in vision

23:03

language such as text, pass it into a model, and then that spits

23:06

out actions for your robot to do.

23:12

So we're going to take a look at all of these three separate

23:16

components and the tools that NVIDIA has to support this.

23:21

NVIDIA Isaac Lab.

23:23

So Isaac Lab is our open source modular framework for robot

23:26

learning and policy training.

23:28

Here you can bring in your own robot assets, any kind

23:32

of embodiment that you're looking into, objects that

23:35

you're trying to manipulate.

23:37

Set that up inside of a simulated environment, have a tele-operated

23:41

device if that is of your choice, and you can do reinforcement

23:45

learning, imitation learning, you can look at certain motion

23:48

planners that will eventually spit out these ONIX files that you

23:52

can run on any environment, whether that be an autonomous mobile robot,

23:57

a manipulator, or a humanoid.

24:00

Now, this GTC, we've released Isaac Lab 3.0.

24:03

And let's take a look at the architecture a little

24:07

bit more in depth.

24:08

The main points that I really want to highlight here are

24:10

kind of in that bottom, towards that bottom layer.

24:14

So Isaac Lab in the past, we had only kind of been able

24:17

to support physics as our renderer and our physics engine.

24:23

But now we have kind of modularized this and decoupled this so you

24:27

can actually swap in instead of if you're if you want to use something

24:32

like Newton, for example, I'll talk a little bit more about all the

24:35

benefits that Newton has to offer.

24:37

But if you're looking to use Newton as your physics engine

24:40

and you can totally swap that out instead of maybe if you

24:44

want it instead of physics.

24:46

This is also the same thing for rendering.

24:48

We have a Warp Renderer that we also support RTX rendering.

24:53

Visualizers, same thing.

24:55

And what's the real kind of advantage here?

24:58

Well, let's take a look at some of our visualizers.

25:02

So at the bottom right-hand corner, you see the Kit Visualizer.

25:05

This is something that if you've played around with

25:07

Isaac Sim before, Isaac Lab, this is something that you're

25:10

probably familiar with.

25:11

You've seen this UI before.

25:13

But we've introduced a new visualizer, if you

25:16

can see it at the top.

25:18

There's also this rerun visualizer available on the web.

25:21

But this allows for you to check your RL training, for

25:25

example, as it's progressing.

25:27

And these lightweight visualizations allow for your

25:30

compute to be utilized for heavier portions of your training pipeline.

25:35

Renders.

25:37

There have been a lot of improvements on the RTX renderer

25:39

side, including RTX Fast Depth.

25:44

That's one of our new features that we've introduced into

25:47

Kit, and we're able to render about two to three times faster than what

25:52

we've previously been able to do.

25:54

We've also included some RGB modes.

25:57

Albedo will retrieve that RGB data of just the colors of your robot

26:02

without any additional kind of like shadows or lighting effects.

26:05

But if you are looking into kind of making sure that you have

26:09

that direct lighting support, well, there's the RTX simple shading.

26:14

We have that available.

26:16

Both of these are a little bit more simplified, lightweight rendering

26:19

options that will allow you to retrieve that RGB information

26:24

from that RTX renderer and they'll run much faster now than what

26:28

they have previously done before.

26:31

Another renderer that we want to highlight is this Newton

26:33

Warp tile rendering.

26:35

So it's another lightweight option, and it's much faster than RTX.

26:40

We've seen very promising numbers from the Mojoco side,

26:44

where we've seen upwards towards like one million FPS.

26:49

And this is a really good way for us to be able to trade

26:52

off the performance versus quality in terms of rendering.

26:56

And again, we really want to highlight that this RTX renderer

27:00

can be used with a PhysX physics engine, Newton physics engine,

27:07

this is all interchangeable.

27:12

So you've heard me talk about Newton.

27:14

Pretty sure you guys have heard about it as well.

27:17

So Newton is our open-source accessible physics

27:20

engine for robotics.

27:21

It was a collaboration with Disney Research and Google DeepMind.

27:26

And a couple of the really kind of key highlights that I want to

27:30

make here are there are generalized simulation for mechanical

27:33

linkages, including closed-loop linkages, especially for RL.

27:38

We've included some high-fidelity contact and grip modeling.

27:42

We've enabled support for deformables.

27:45

If you're looking into kind of with cables, clots, deformable

27:49

parts, we have a couple of solvers that I'll go into in depth more.

27:55

And then GPU accelerated for robot learning at scale.

27:59

So Newton is a physics engine that we're really highlighting here.

28:04

In order to make sure that you are able to do all your

28:08

robot learning at a much adequate and really utilize those GPUs.

28:16

So let's look a little bit more on the Newton architecture.

28:20

So, again, Newton, it is modular as well, and I want you to

28:24

focus on that middle part of the graph that you see there.

28:28

So these solvers, Majoko, the Majoko solver is one of

28:34

the main solvers that many of you have heard of, but we also support

28:38

Camino solvers, deformable solvers, canonical solvers, These solvers

28:47

Perfect.

28:48

These solvers allow us to kind of, especially what we've seen now,

28:54

for robotics we really want to make sure that we're not only working

28:58

with just rigid objects, we have so many articulations going on,

29:04

a lot of materials with the objects that we are interacting with.

29:08

The Camino solver that you see handles complex and intricate

29:14

closed chain mechanisms.

29:19

And the vertex blocks descent solver, the VDB solver that

29:23

you see at the bottom left-hand corner, that one allows you

29:27

to handle those, like, linear deformables such as cables.

29:31

And any kind of like rubber parts.

29:34

You have the implicit NPM solver on your right, and

29:37

that handles particle simulations that are applicable to rough

29:41

terrain and locomotion scenarios that you might use.

29:48

And so, having a really good physics engine has been super,

29:53

super helpful in reducing what we call the sim-to-real gap.

29:57

So, how can we really address that?

30:00

Well, here I'm showing an example of a locomotion policy

30:03

on a quadruped, the Animal D, that was trained with Newton at first.

30:08

Then it was, and then we used physics to do

30:13

the sim-to-sim transfer.

30:15

And then from there, we did the animal Newton to real.

30:19

And you can see that this works.

30:23

This works. It transfers.

30:25

So just really want to make those callouts and that highlight.

30:29

So the Isaac Lab is a lightweight, modular,

30:32

multi-physics training framework.

30:34

It integrates Newton's physics engines, physics, and Omniverse to

30:39

really scale up your RL workflows.

30:42

And then you have these fast, perceptive learning

30:45

for DGX class GPUs.

30:50

So, talked a little bit more on the updates that we had for Isaac Lab.

30:56

In Isaac Lab, now I want to kind of switch on to something else.

31:01

I mentioned VLAs at the beginning.

31:05

So this GTC, we've released Isaac, GROOT, and 1.7.

31:09

These are VLA where, let me just kind of go over the architecture

31:17

of how VLA's work.

31:19

Again, you have an image, for example, as an input text.

31:25

And then we pass it on to a vision language model.

31:29

So since the last release, starting off with GROOT 1.6, we've actually

31:35

changed that vision language model to use a Cosmos 2 billion

31:40

variant model in the back end.

31:43

Then passed on to a diffusion transformer model that will

31:46

later give us some action and tokens for our robot to use.

31:49

This model has a commercial license, fully customizable,

31:54

and supports different embodiments.

31:58

Let's take a look at some of the embodiments that GROOT has

32:01

been able to have been deployed on.

32:04

So we see here Unitree G1 placing a box on its shelf,

32:08

and we see here the locomotion and manipulation tasks as a whole.

32:12

We also see agibot, bussing a table, yam, folding a shirt.

32:19

So we've included these pre-trained weights for zero-shot evaluation,

32:24

but fine-tuning the model is really beneficial when you're deploying

32:27

to a specific embodiment or task.

32:31

Just to really call out those highlights on what GROOT and 1.7

32:35

has, again, factory floor ready, has a commercial license, which

32:39

enables production deployments across material handling,

32:42

packaging and inspection, reasoning for multi-step tasks, and then

32:50

expanded dexterous manipulation.

32:56

So, in theory, we've used Isaac Lab to kind of train a policy.

33:03

GROOT is a sample model that we can use.

33:08

But now let's say that I have a policy.

33:11

Now, how do I really make sure that that policy, those

33:14

actions that I'm getting from the trained policy, how do

33:18

How do I make sure that

33:20

My robot is able to perform those.

33:23

So, in that, we have a controller.

33:27

So, Sonic is a humanoid whole body controller.

33:30

And what is really great to call out here is that we've had

33:34

controllers in the past that were just strictly kind of like just

33:38

locomotion-based, or we've just had manipulation-based controllers.

33:42

But Sonic is something that really kind of, it's a single versatile

33:46

A policy that handles all of these diverse tasks, such as locomotion,

33:50

such as jumping, manipulation.

33:52

Rather than having all of these separate controllers,

33:55

Sonic is something that is able to combine all of these together.

34:00

Let's look at a couple of examples.

34:02

So with Sonic, we are able to really perform and see

34:06

this really fluid-like motion from a lot of what you would

34:10

consider human-like behaviors.

34:12

These are a couple of examples here.

34:14

There's a full body VR tracking.

34:17

Example in the bottom middle of your screen there that

34:22

captures an operator's complete body motion and you can see

34:26

how swift and easy that is.

34:28

There's a couple of music ones and being able to, you

34:33

know, kind of like any text that you want to provide it, it can

34:40

perform similar actions like that.

34:43

So, that connects with your policy.

34:47

And so, you might be looking at me and you're, okay,

34:51

so now I feel good.

34:52

Let me go and deploy this on a physical robot, right?

34:58

We don't want to incur a lot of damage and I'm pretty sure

35:02

you don't want to have to tell your manager I broke a motor.

35:06

Can we place an order for another one?

35:08

Well, that's why we introduce policy evaluation.

35:12

So once you've trained a policy, before going on to physical

35:18

deployment onto a real robot, we really want to make sure

35:22

does my policy work and how well does it work?

35:28

We've introduced something called Isaac Lab Arena.

35:32

And where does this sit?

35:37

So Isaac Lab Arena.

35:41

So Isaac Lab Arena is our open-source framework

35:44

for large-scale policy evaluation and simulation.

35:47

It extends from Isaac Lab.

35:50

So the main things is we're able to do task curation and

35:54

policy evaluation.

35:55

Those are the two big things that Isaac Lab Arena has to offer.

35:59

So if I set up my environment inside of Isaac Lab, Isaac Lab

36:05

Arena, it's an extension of that.

36:09

So I can actually create as many different environments

36:13

that I want, switch my, for example, if I had a manipulation

36:20

task to pick up an object, right, I can randomize the starting and

36:25

ending position, starting positions of this object, and place it in as

36:29

many different kind of environments that I want, and I want to

36:33

see if my policy actually works.

36:35

So Isaac Labarino allows you to really kind of scale those

36:39

tasks from one to end tasks.

36:40

You're able to set metrics as to what is successful for you.

36:46

If I'm picking up my water bottle and it kind of like

36:50

falls along the way or it slips out of the robot hand, is that to

36:54

you considered successful or not?

36:57

And so Isaac Laburina allows you to kind of question those things.

37:02

I want to show here a couple of examples that we've, that

37:06

us and our partners have used.

37:07

So I will kind of start from the top left here and then go around.

37:13

So in the top left, we have a GROOT industrial benchmark.

37:17

This is evaluating a GROOT policy.

37:20

Then we have, in the bottom there, we have RoboTwin benchmark

37:25

evaluating an ACT VLA policy.

37:27

At the top right, we have a NIST board benchmark evaluating

37:31

a reinforcement learning policy.

37:33

And at the bottom left, we have PI0, 0.5 VLA policy.

37:40

And so I mentioned not only VLA models, but also reinforcement

37:45

learning as well.

37:46

You can bring any VLA model, any policy that you've trained

37:49

into Isaac Lab Arena, and you can test it and evaluate it and see

37:53

how good it performs right before going on to that deployment stage.

37:58

So this is a really, really great way for you to make

38:02

sure and feel confident before going on to that step.

38:05

Just to kind of wrap up all of the real kind of key values

38:11

that ARENA has to offer, we're able to accelerate this task chaining

38:15

skill as well that I will mention a little bit more about now.

38:20

So in the past, we've seen kind of just a pick and place

38:23

or if a closed door, for example.

38:26

ARENA also allows you to do task chaining, which I think

38:29

is really, really important as we're now starting to look into

38:34

Robots that are starting to, we want them to do more complex tasks.

38:39

And so Arena allows you to play with that as well.

38:42

You're able to do scene authoring with relational object placement.

38:46

The real benefit in being able to scale up your evaluation

38:50

benchmarks is to be able to swap objects in and out with ease.

38:54

So Arena actually has, it's built.

38:58

It has text command instructions.

39:01

So if I want to say place my object on top of a microwave,

39:06

I'll place it next to.

39:08

All of these commands in this format allows us to quite

39:13

easily switch objects in and out.

39:15

Eventually, this is something that we are looking forward to, like,

39:18

an AGENTIC kind of, like, workflow.

39:22

But it's definitely something that we're headed towards.

39:26

Diversity in objects and just really being able to

39:30

streamline our workflows for training and evaluation.

39:33

So this really allows us to kind of like close the loop

39:36

and we're training, then we're going to evaluate and then

39:39

train again until we get a policy that we're really happy

39:43

with and we think is successful.

39:50

So that covers kind of what we have for training.

39:54

We talked a little bit about Isaac Lab, we talked about

39:56

GROOT, we talked about Newton, we talked about ARENA.

40:01

So theoretically, after you're done in this step, you do have a model,

40:05

a policy that you feel confident.

40:09

Let's take a look at what tools we need, but also how that looks

40:17

like to build a humanoid robot.

40:19

So I'm going to start top down.

40:23

So here we're starting off with our high-level reasoning, right?

40:26

So these are some VLMs or some reasoning models.

40:30

From what Akul has talked about, Cosmos' reason falls

40:35

under this umbrella.

40:38

Then you have perception and planning, so here are where

40:41

your VLAs come in, whether you have any navigation, locomotion

40:47

kind of models, manipulation, you're trying to say what's the

40:52

object, what's the goal that I'm trying to do and how I can do that.

40:56

Then you have your real control framework.

40:58

So here are where my whole body controllers live.

41:01

I talked about Sonic.

41:02

That's where you would find that in this layer.

41:05

And then towards the very bottom, you'll have this

41:08

hardware abstraction layer where your sensors are living

41:11

at, your cameras, your LIDARs, IMUs, whatever you're looking

41:15

into, that's where that would live in this stack.

41:18

But all of this has to run somewhere.

41:21

I need a computer on the edge.

41:25

I need a computer that will live inside of my humanoid robot.

41:27

There is no way my humanoid robot will carry my huge workstation

41:32

and perform all of these tasks.

41:34

So, to do that we have Jetson Thor.

41:38

So, this is built on NVIDIA's Blackwell architecture.

41:42

Jetson Thor has 128 gigabytes of memory, and it introduces

41:47

native FP4 quantization, so it has a transformer engine

41:53

that's able to switch between FP4 and FP8, which will allow

41:57

for that optimal performance.

42:00

Jetson Thor also introduces multi-instance GPU.

42:03

So your GPU can actually be partitioned into isolated

42:07

instances with dedicated resources, meaning that whenever you

42:11

have workloads that are running either really high or really low

42:15

intense tasks, it's able to kind of like handle those situations.

42:21

Let's look a little bit more on Jetson Jetpack 7.

42:26

Well, Jetson, NVIDIA Jetpack is the official software stack of NVIDIA

42:31

Jetson, and it's able to provide a comprehensive suite of tools.

42:35

We're able to run a whole...

42:38

A variation of generative AI models from VLAs like GROOT to some

42:46

really popular LLMs, VLMs as well.

42:50

We also include a lot of support for NVIDIA Isaac, this is

42:56

where Isaac Lab would sit, also Metropolis for visual AGENTIC AI

43:02

and Holoscan for sensor processing.

43:07

Let's look at the different kind of performance metrics

43:10

here so many of you guys have probably heard of our previous

43:13

version which was Jetson Orin and if we like let's just

43:18

looking as to how Thor will compare

43:23

If I had a Jetson Orin running one VLM, for example, for

43:28

just real-time sensor processing, another one for an LLM, just

43:33

a summarization tool, and then a final Orin to actually

43:38

do the robot, like, actuations, I would need three separate Orins.

43:43

But in this case, Jetson Thor is powerful enough that I can reduce

43:47

all of those three and suggest one.

43:50

So that really kind of like just shows you the power of Jetson Orin

43:57

and how much it's able to handle.

44:04

Aside from just compute and being able to run any kind

44:08

of model that I want, I also want to make sure that it's able to

44:12

integrate to a lot of the sensors that I will be working with,

44:16

so cameras, IMUs, actuators, I need all of that to be connected in.

44:21

So, with NVIDIA's Holoscan Sensor Bridge, you can seamlessly connect

44:26

all of these sensors over Ethernet, regardless of the modality.

44:30

And using their camera over Ethernet technology available

44:34

in Jetson Thor, it's able to stream this data directly into GPU memory,

44:39

dramatically reducing that latency and minimizing that CPU overhead.

44:49

Another thing that I really want to highlight is our safety,

44:54

our robotic safety platform.

44:55

So Thor IGX supports both inside-out and using on-board

45:00

sensors and outside-in safety using infrastructure sensors,

45:04

but at its heart, there's a dedicated functional safety

45:07

island that provides an independent safety processor that isolates

45:12

all of these critical workflows.

45:14

It's designed to mean ISO 262 and

45:20

IEC 61508, so these are all just kind of, and whenever we're looking

45:25

into our sensors and if we're noticing that something's going

45:28

wrong, we do support, we do have this safety platform that's able to

45:33

kind of like kick in when needed.

45:38

I've talked a little bit about Thor.

45:42

I've talked how our hardware is, or how this is the computer

45:48

on the edge for humanoid robots.

45:51

Now I want to take a step back.

45:54

I wanted to introduce NVIDIA OSMO.

45:57

So NVIDIA OSMO really isn't a hardware.

46:01

Um, but I think I had it felt that it kind of

46:08

So NVIDIA OSMO is our open source orchestration platform

46:11

for building, testing, and validating physical AI.

46:15

So building robots, as you can imagine, you probably

46:19

have a model running on a cloud or you have your work

46:24

is like split in between laptops, your workstation, the cloud.

46:28

How do you orchestrate this so that it all just kind of

46:31

like works together?

46:33

Well, we've released, so OSMO is built to kind of

46:39

like help for that.

46:40

So developers can really simply define their workflow using a

46:44

YAML file that eliminates the need for all of this complex scripting.

46:49

So you're able to kind of like...

46:50

If you need allocation for data generation, then you

46:56

need to go on to model training.

46:57

You need something that can orchestrate things in a really

47:00

fast and fluid manner.

47:01

And so NVIDIA OSMO is the tool that can help you for that.

47:06

We've announced the physical AI data factory blueprint.

47:10

And here, I kind of want to call out, so in order for

47:14

this to happen, we've talked about these tools, so I want

47:19

to do data acquisition and data generation, right?

47:23

I want to collect my data.

47:24

Then I want to augment my data.

47:26

Then I want to pass that on to training.

47:29

And how do I make sure that all of this is kind of like

47:32

working together?

47:33

Well, OSMO is what really kind of like helps you

47:36

orchestrate those tasks.

47:39

And so we've partnered with Microsoft, Azure,

47:43

and NEBIUS to host this.

47:47

And we have a couple of partners like Field AI, Hexagon, Milestone,

47:51

Skilled AI, and Teradyne that are working and are using blueprints

47:57

like these in their workspaces.

48:01

Cool.

48:03

So, now that we've gone through data generation, policy training,

48:10

and deployment, I wanted to take a look back at this slide,

48:16

because now I will show you an example that we've done

48:19

end-to-end that incorporates everything working together.

48:23

You've most likely... Yes, This is the video.

48:33

So here we see

48:38

This is a, we've trained end-to-end, this is in our

48:43

NVIDIA headquarters, we see here, so she's putting

48:49

a task, bring me the healthiest snack, the robot then kind

48:52

of takes in that snack, takes in that command, there's Cosmos Reason

48:57

running on in the background, analyzes what it has to do, there's

49:02

a navigation policy that has been trained inside of simulation,

49:06

To make sure that it's handling,

49:10

to know where it's going, how does it focus on following a direct

49:16

path, reaching its final obstacle.

49:21

And so it's getting closer to that table, from that egocentric

49:25

view it's able to identify and it's thinking.

49:30

Okay, there's a couple of items here, but the middle has

49:34

the apples, which are good for you.

49:37

And the right has cookies, which is also not the best.

49:41

But then it goes ahead and makes that decision to then

49:43

choose that apple.

49:45

And now, it's thinking, okay, now I have to go back and

49:48

complete my cycle.

49:49

Let me go return this apple back to the person that has kind of

49:55

commanded this instruction for me.

49:57

So there's a navigation kind of like pipeline going on there.

50:01

We have a couple of egocentric, in order for it to identify

50:05

where it's at in space, we've We've also kind of worked

50:08

with a couple of vision-centric mapping and localization tools.

50:13

But again all of this is working together.

50:16

We have our whole body controller working, making sure that

50:19

it's stable and then it's finally able to get there.

50:28

Thank you.

50:30

Awesome.

50:31

Yeah. So this is a really great way as to how we're

50:34

able to kind of really combine all of the tools that we've seen

50:38

today, how they all kind of come together to produce this example.

50:45

So, again, I'm providing this kind of, like, reference workflow

50:48

as to I get this question all the time where it's how

50:51

do you break apart a robot?

50:53

How do I go ahead and do something like this?

50:56

It's really hard to wrap my head around.

50:58

Well, I've included this kind of, like, high-level overview

51:02

that includes kind of data generation, augmentation,

51:06

training, and evaluation, and then, finally, deployment.

51:09

So hopefully this helps.

51:11

But this concludes our talk where we're covering the whole

51:17

stack, everything that NVIDIA has to offer, and I know this

51:21

is towards the end of GTC, so I am really grateful to see everyone

51:25

here, really excited to learn and hopefully build more robots.

51:32

But with that, thank you, and we are opening up the

51:34

floor for any questions.

51:37

Oh, cool. Yeah.

51:40

Thank you, Edith and Nakul.

51:42

That was a really great, wonderful deep dive on how

51:46

to build an end-to-end pipeline for humanoid robots.

51:49

Now for questions, we have two mics.

51:52

You can create a line and then go ahead and ask your questions.

51:57

Very nice presentation.

51:58

Thank you very much. It's really informative.

52:00

So my question is, how do you evaluate the data

52:03

generated by Cosmos?

52:05

Can you repeat the question again?

52:06

How do you evaluate the data generated by Cosmos?

52:10

How do we evaluate the data?

52:14

Evaluate as in whether it was good enough or good or

52:18

bad or whatever it is.

52:19

So you can use Cosmos Reason for that.

52:21

So Reason can look at the generated videos from Cosmos

52:24

Predict or Cosmos Transfer and look for physical inaccuracies or

52:33

any other type of hallucinations.

52:35

And you can use Reason for either rejecting or accepting.

52:38

So we have a thing called Cosmos Evaluator, which was

52:41

built using Reason, which does it.

52:43

It's already on GitHub.

52:44

You can use it for this task.

52:48

Awesome. Hey, my name is Jan from Andromeda.

52:51

Thank you for the presentation, very informative.

52:54

What I wanted to ask, which is something that's been on

52:57

the floor, on the convention floor for quite a bit, but no one's

53:01

given me like a good answer yet, which is the HoloScan technology.

53:06

Obviously, it's quite new, just reference designs.

53:08

How many people is actually using that at the moment?

53:13

What I see is there's a gap of, like, ready-to-use boards

53:18

for that technology.

53:19

It's like, well, there's the, you know, lattice ICs, you

53:23

know, et cetera, but there's no, there's not enough boards

53:26

that have been built up yet.

53:28

So I'm just wondering, like, you know, from your side,

53:30

how many people are using it?

53:32

Like, what have their experiences been, et cetera?

53:36

Yeah, I don't know too much about the partners that we're

53:40

actually working with, but I'm more than happy to kind

53:43

of, like, dive into this a little bit more and kind of, like,

53:46

get you the answers that you need.

53:48

And towards the end, you said, can you repeat that last part?

53:51

Like, I don't know if there's...

53:54

enough off-the-shelf or ready-to-use commercial

53:58

boards for those adapters for different kinds of things.

54:01

I think there's a lot of custom designs being made by the

54:05

IC makers to show off designs, but for actual commercial ready-to-use,

54:11

like COTS kind of grade.

54:15

Yeah, especially for a lot of custom designs, we do work

54:18

with a lot of our partners.

54:19

So I'll have to get in touch with the right person.

54:23

I'll get the right person in touch with you so that

54:25

we can talk about that further.

54:26

Sounds good.

54:30

Hey. Thanks for the great talk.

54:32

I had a question about Cosmos and kind of the roadmap for that.

54:37

It seems that currently it's mostly focused on, at least for Cosmos

54:43

transfer, focused on RGB data.

54:47

Do you guys foresee that eventually evolving to LIDAR and depth and

54:53

time of flight or other modalities?

54:55

For robotics related stuff?

54:59

I'm not fully sure how transfer itself will evolve, but I

55:03

think if you have attended Mingyu's talk, he did talk about

55:06

Cosmos 3, which will be coming out later this year, which will be

55:09

like a single unified omni model, which will be having all these

55:13

different capabilities of predict, transfer, and reason together.

55:18

I still don't know if it will be able to generate other

55:21

modalities, but you never know, I think as we get closer to that

55:28

we will be able to answer that.

55:29

Okay, yeah, I wasn't able to attend that talk.

55:31

Is there more information about Cosmos 3 online already?

55:34

No, like I said, it's later this year, so we just are

55:37

announcing the name and the general architecture for it.

55:44

He is planning to release in Q3 sometime, and for Mingyu's

55:50

talk, it will be available within 72 hours, that happened yesterday,

55:54

so yeah, you can watch it online.

56:00

Hi. Great session, guys.

56:01

Very, very informative.

56:03

My question is a little bit more about real-world data.

56:06

We were looking at a lot of vendors kind of providing, say,

56:10

egocentric data or teleop data.

56:12

The first part of the question is, where do you think egocentric data

56:16

is, you know, useful, and where does teleop become more required?

56:20

And the second part of the question is, can Cosmos Reason

56:23

be useful in, you know, checking quality of the data that's

56:27

collected from the real world?

56:29

Checking quality of the data that's collected from the real world.

56:33

Okay, that's interesting.

56:34

So to answer the first part of your video, egocentric

56:36

data is very important.

56:37

Like as I explained in the Dream Dojo part of the presentation,

56:42

we trained the Cosmos Predict model with 44,000 hours of

56:47

video, and then with smaller amounts of teleoperation data sets,

56:51

like we did the post-training after that, and it was able to generalize

56:56

to completely Really different objects and different environments.

57:00

And that ability kind of comes from the large amounts of training that

57:04

we did with the egocentric videos.

57:08

We also have some other work that's happening that already

57:12

got published called EgoScale, where we also showed that there is

57:16

like a clear scaling law, where the more egocentric data that you're

57:21

adding while you're training a VLA, the better the performance gets.

57:24

Now, these are all research works, so all the learnings that we have

57:28

from research, all these different research works, will kind of come

57:31

together for the future releases of GROOT and Cosmos models, which will

57:35

be how we'll be able to use them.

57:37

So just to follow up on that, so would you add a lot more teleop

57:41

data to the pre-training mixture?

57:43

Or would it be mostly for post-training?

57:46

For the GROOT kind of models, or like the Cosmos models?

57:49

Yeah, any kind of foundation models.

57:51

So, Cosmos, again, is a world foundation model, so we don't want

57:53

it to be tied to a particular robot embodiment or anything like that.

57:57

So, teleoperation data from a robot will definitely be

58:00

a post-training thing, right?

58:02

And I just wanted to add to what Akul said, we also just

58:07

launched Cosmos Evaluator, which is based out of Cosmos Reason, it's

58:11

already published on GitHub, you can feed any domain specific data

58:18

to that and it has inbuilt checkers for robotics, AV, visual AI agents.

58:23

So you can customize it for your own use case, and it

58:27

also integrates really well with some of the top coding

58:32

agents today, like Cloud Code and Cursor, whatever you like.

58:38

So yeah, I would highly recommend trying that out.

58:41

And you can also build your customized checkers if you

58:44

do not want to use what's already in there.

58:47

Even for real-world data?

58:48

Yes. Awesome. Thank you so much.

58:50

Thank you.

58:52

Hi, my name is Melvin, my name is Melvin Valdez, I work for

58:58

a small company named Certainty, we do data layer encryption,

59:05

self-protecting data, and in that field, I was wondering if

59:15

There is that kind of encryption in all this where the data

59:22

being sent and all that is

59:26

If y'all have looked into the encryption of that,

59:28

keeping the data safe and... The encryption of the data?

59:32

I didn't understand. Or, like, encrypting the data so it's

59:35

encrypted during the use, during the transfer and all that.

59:38

If that's something NVIDIA's looking into.

59:42

Like data formats?

59:43

I think the question is that, while training, is there an

59:47

ability to encrypt data?

59:49

Is that the question? Yeah, if that's something that...

59:51

I'm not super aware of this.

59:56

Which allows you to be able to encrypt things and decrypt

59:59

it at the hardware level, which is probably how you

01:00:03

do it, but like, yeah, I mostly work on the software end of

01:00:06

things, so this would be like more of a hardware level question.

01:00:08

Okay, thank you.

01:00:11

Awesome.

01:00:12

Sorry, I'm back again.

01:00:15

I got two more questions. One, you mentioned the Safety Island.

01:00:18

Is that just on the 4 or is that on the, or in AGX as well?

01:00:22

That's on the 4, yeah.

01:00:24

Okay, cool.

01:00:26

And then, I was in a few talks about AlphaMail, like the

01:00:30

self-driving car aspect, and, you know, talked about how

01:00:34

the system was built and...

01:00:36

It's like a reason-action model.

01:00:39

I guess GROOT is similar, but I guess for them, they

01:00:44

were able to do some safety certification on that, as

01:00:48

well as the rest of the stack.

01:00:51

If we were to employ some of the SDKs that are on the

01:00:55

Jetson-Ross side, are there any Safety certs, et cetera, pathways.

01:01:03

So the question is, it's like if with the Alpha Mayo stack,

01:01:10

their promise was that it can be safety certified to,

01:01:16

you know, whatever the standard is for employing this kind of stack is

01:01:23

their safety certification pathway.

01:01:26

So for robotics right now, we don't have that at that level.

01:01:30

So I think what you're trying to mention is that in the

01:01:32

AV side of things, we have different levels of certifications

01:01:35

for models and functional safety and different things.

01:01:38

In the robotics side with IGX, we have functional safety,

01:01:40

but we don't have anything about that at this point.

01:01:43

Yeah, but there is a kind of draft safety standards

01:01:48

coming out, I believe.

01:01:51

But we can talk later.

01:01:55

Hi, I'm Fabian, I'm from a company called Qualia.

01:01:58

We do VLA infrastructure.

01:02:00

I'm wondering, with GROOT 1.7, you mentioned it's commercially

01:02:03

available, what's the best way to get started with it?

01:02:06

I know it's not available today, sort of, everywhere,

01:02:10

but what's the plan there?

01:02:12

Yeah, I'm re- Sorry, can you repeat your question?

01:02:14

With GROOT 1.7, what's the best way to use it today?

01:02:21

I believe it is coming out soon, but we also have our

01:02:25

amazing product marketing person for GROOT, Kalyan,

01:02:30

and you can catch him up and talk more about GROOT if you want.

01:02:34

Yeah, he'll get you with the right answer for that.

01:02:40

Since GROOT 1.7 is our first model where we're going with a commercial

01:02:45

license, all the previous versions haven't had that commercial license

01:02:49

support, so we'll definitely go ahead and talk to Kalyan for that.

01:02:55

Hi, thank you for the talk.

01:02:57

It was amazing. But I have one question.

01:02:59

When you're deploying the entire Cosmos and Isaac stack

01:03:03

in real humanoid environments, what's the first thing to

01:03:06

break in the sim-to-real gap?

01:03:09

Like, what's the first thing that stops working?

01:03:13

Is it control, perception, inference, or which part of

01:03:18

it stops working first?

01:03:21

There are multiple elements to this.

01:03:26

For the Cosmos part, at this point, I think the biggest challenge

01:03:30

is that the models are not small enough or fast enough, because

01:03:35

of the complexity of the Cosmos models, that they can't be run

01:03:39

on an edge device at this point.

01:03:41

So all the experiments that we saw, we were running around

01:03:43

more powerful compute.

01:03:45

Future versions of GROOT, and like I said, these are all

01:03:48

research works and we were learning different things from them.

01:03:51

So the learnings from that will go into the future versions

01:03:54

of GROOT and Cosmos.

01:03:56

And there you will be able to use those things in edge

01:03:58

devices and things like that.

01:04:00

That's one part.

01:04:01

The second part, for the simulation-trained policies

01:04:07

or VLA policies, the things that, honestly, there's a lot of things

01:04:11

that can break, there's no one single thing that I, or the first

01:04:15

thing that will always happen.

01:04:18

One common thing that generally happens in simulation is that

01:04:23

the modeling of the actuators, the modeling of the different

01:04:28

parts of the humanoid and things like that would be

01:04:30

maybe wrong in the simulation compared to the real world.

01:04:33

And that's like one of the first things that you immediately

01:04:35

see if you try to deploy it on a real robot.

01:04:37

So usually what we recommend people is that we do sim-to-sim

01:04:40

transfer and see if the policy hasn't overfitted

01:04:45

to a particular simulator.

01:04:47

So we don't want the policy to learn how a particular

01:04:52

simulator operates a particular joint or things like that.

01:04:55

We want it to be above that and then test it

01:04:57

on a different simulator, like from Newton to physics.

01:05:01

And then if that transfer works, it's very likely that

01:05:04

it will work in a real robot.

01:05:07

We'll just take a couple of more questions from people

01:05:09

who are already in line because we want to give them a break, they are

01:05:13

talking non-stop for nine minutes, but you can cast them in the lobby.

01:05:18

Hey, my name is Sanjeev.

01:05:20

Thank you. The presentation was great.

01:05:22

I have a question. You mentioned that for multi-agent workflows,

01:05:26

Jetson Thor would handle it the same amount as three Jetson Orin.

01:05:31

My question is, in the last video we saw where the G1

01:05:33

was retrieving the apple, was everything running on the edge?

01:05:39

Because to the best of my knowledge, the G1 has only

01:05:43

two Jetson O-ring, so this was able to run all three things like

01:05:48

VLM, LLMs, and VLAs on the edge.

01:05:52

Yeah, we had GROOT running on the edge on that, we had

01:05:55

Cosmo, since the VLM for, that one was GROOT and 1.6 also had

01:06:05

Cosmos Reason as a VLM, so it, they wasn't, it was running all of those

01:06:10

things kind of like on the edge, on the... Yeah, I think there's like

01:06:14

two parts to that, I think the...

01:06:16

The topmost layer of VLM doesn't necessarily need to run on the

01:06:20

edge, that can be outside, and can be just an inference because you

01:06:24

don't have to run it every single, like every 5 milliseconds or

01:06:28

something like that, but GROOT also uses Cosmos Jetson as a VLM within

01:06:35

its architecture that was running.

01:06:39

Like a backpack with Thor that you can attach to a Unitree robot right

01:06:43

now and be able to get the full Thor ability even in the older G1s.

01:06:49

Okay, thanks.

01:06:52

Hello, thanks for the talk.

01:06:53

My question is, how can we bridge the gap between human

01:06:57

data and robotic data?

01:06:58

Since a world model needs to predict the next frame,

01:07:02

if that next frame contains gaps inherent in human data,

01:07:06

how can a robot effectively leverage that knowledge?

01:07:09

Thank you. Yeah, I think, especially for our VLAs and...

01:07:16

We train model, like our GROOT model is trained on a wide

01:07:19

variety of data, right?

01:07:21

This is tele-operated data, this is egocentric data, this is data from

01:07:25

YouTube that's available online.

01:07:28

And so having all of this data available, we're able

01:07:35

to pre-train a model that will at least give us a good understanding

01:07:39

of how the physical world works.

01:07:41

Now when you're coming into kind of like post-training

01:07:44

for like a specific embodiment or if you wanted to learn how

01:07:47

to do a specific task, then that's when you would want to kind of

01:07:51

like collect data for that reason.

01:07:54

And so human data still really is like egocentric data still

01:07:58

is really useful.

01:08:00

You're able to kind of like map your kind of like arms,

01:08:05

your hands into kind of like a robot embodiment.

01:08:10

I think all all data is is really important and it's

01:08:13

really useful um so there isn't kind of just like oh would I

01:08:18

prefer one versus the other I think they're all they're they they can

01:08:22

all be really useful and different and it could be in pre-training

01:08:26

and post-training um but yeah.

01:08:31

Okay, thank you.

01:08:34

Can I proceed with the question? Yeah, sure.

01:08:35

Oh, perfect.

01:08:37

First of all, thank you for the presentation.

01:08:40

When training GROOT with egocentric data, what kind of pre-processing

01:08:44

steps do you recommend to do to achieve a successful training?

01:08:47

What kind of information is basically essential to

01:08:50

extract from the videos?

01:08:52

Yeah, so kind of in that, like I said, kind of like

01:08:57

tracking, not only like your arms, your joints, to kind

01:09:01

of like a robot embodiment, I think that that would be the most crucial

01:09:05

steps, you really want to make sure that you're kind of like being

01:09:09

able to track all of those joints.

01:09:12

The egocentric data, I think, is also, it's really interesting

01:09:16

because sometimes even like fixing your, some examples that

01:09:23

we've had are a fixed camera versus a kind of a camera that can move.

01:09:28

And so a camera that can move kind of is able to capture

01:09:32

a much wider field of view.

01:09:35

And in these cases, we've kind of had to learn as to

01:09:40

kind of new noise gets introduced in our models and new kind

01:09:44

of if we have distractors in the background that you didn't see the

01:09:48

model can kind of get distracted.

01:09:51

And so that kind of allows us to kind of learn a little

01:09:56

bit more as to how we can block out those noises.

01:09:59

But yeah, that's kind of.

01:10:01

I would make sure to focus on being able to map your vision and the

01:10:10

information that you're seeing and convert that to a physical robot.

01:10:15

A really fast continuation from the previous question

01:10:18

that you answered.

01:10:19

Does NVIDIA offer any kind of models or services to extract

01:10:23

that kind of information?

01:10:25

From egocentric data?

01:10:27

For example, as you mentioned, the joint positions or segmentations.

01:10:32

From egocentric data.

01:10:34

So, if you had watched the Isaac Teleop presentation,

01:10:39

one of them, they had two of them, they actually announced something

01:10:42

called Ego4Robo, which is something that will come later this year,

01:10:47

which will allow you to do that.

01:10:48

Okay.

01:10:51

Hello, I just had two questions.

01:10:54

One is like how much cost and compute power is needed

01:10:58

to let's say train these world models and right now it's just

01:11:02

using two modalities like RGB and then teleop at the end, but if I

01:11:07

wanted to add one more modality and do a joint training how much cost

01:11:14

or how do you envision doing this?

01:11:16

And the second one is like what do you think would unlock

01:11:19

dexterous manipulation?

01:11:21

Okay, so post-training a world model is a significantly

01:11:25

cheaper thing to do.

01:11:26

You can probably get away with like eight H100s, like

01:11:29

one node of H100s, and depending on how big your dataset is,

01:11:34

a couple of hours of training.

01:11:37

Obviously, creating like an entire world foundation model

01:11:40

from scratch is a...

01:11:42

Like a very expensive operation requiring hundreds

01:11:44

of thousands of GPUs.

01:11:46

So that's the reason why we have this open world foundation

01:11:49

platform, and we are making all these models openly available,

01:11:53

so that you are able to take the pre-training and the work

01:11:57

that we have done, and then post-train it for your specific

01:11:59

physical AI application.

01:12:01

And for dextrous manipulation, what do you think?

01:12:04

Like, just two modalities would be enough, or we need to have

01:12:07

more modalities to have complete, proper dextrous manipulation?

01:12:11

Dextrous manipulation using a Cosmos policy?

01:12:14

Yeah, like in general for humanoids.

01:12:17

I think every example that I showed only had cameras, nothing else.

01:12:20

Oh, okay. Thank you.

01:12:22

Last question. Yeah, thank you for the Q&A.

01:12:25

And I got a question like you said GROOT 1.7 is factory flow ready,

01:12:31

but VLA is the probabilistic model.

01:12:37

So, in a factory event, 0.1 failure is a big deal, so how

01:12:44

do you guarantee the repeatability and is there any something

01:12:48

beyond safety island and handle the task level accuracy?

01:12:54

Okay, I think these are kind of like two separate things.

01:12:57

So kind of like that safety kind of like Island was kind

01:13:01

of like on the Thor side of things.

01:13:02

But for VLA, when I was talking about kind of like factory

01:13:06

ready, those are kind of use cases that we've seen a lot

01:13:09

of our customers deploy GROOT

01:13:12

in factories and warehouses.

01:13:15

So that, when I had that there, that's what I was implying

01:13:18

towards, that this is commercially like available for, and we've

01:13:22

seen it in those popular use cases.

01:13:24

And now that safety island for Thor, like that's, that's separate.

01:13:35

Thank you for the talk.

01:13:36

I have a quick question on the product design.

01:13:38

So, historically, when we have consumer products, they

01:13:43

are mostly manipulated and used by the users of humans.

01:13:47

In this case, it's a robot doing most of these things,

01:13:50

manipulating in the environment and the objects around.

01:13:53

So how does user study design change here?

01:13:57

Does the ecosystem that you guys are developing, does that also

01:14:02

take care of the user-study design?

01:14:05

Or are there any signals that you are already feeding into

01:14:07

this or is this something that your customers take care of?

01:14:11

I didn't fully understand.

01:14:13

What design was that?

01:14:14

It's a user-study design, essentially the product, but

01:14:17

in this case, let's say humanoids, if they have to operate in the real

01:14:22

world, so what is the right design?

01:14:25

Is it usable? Is it reliable?

01:14:28

So is this something that's also being fed by the system?

01:14:32

Essentially because now we are not essentially using the main

01:14:35

driver or the user of the product.

01:14:37

The product is essentially, you know, manipulating them.

01:14:40

I think this is like, this is going to be like an evolving thing,

01:14:44

and this is going to take some time for people to get really right.

01:14:47

I think this is a thing that has to happen at a hardware level.

01:14:50

Like right now, if you take one of the common humanoids

01:14:53

that people are using.

01:14:55

Almost all of them have pinch points, which means that if

01:14:58

you are close to it, and if it's doing an action, you can get

01:15:03

caught while standing next to it.

01:15:06

So these pinch points are there, or most of them are quite hard, and

01:15:10

they are not really that compliant.

01:15:13

Which means that if you try to move its arm or something

01:15:17

like that, it won't be compliant, it won't move with it.

01:15:20

So these are the kind of user design considerations that

01:15:24

you would have to make when you're creating these humanoids

01:15:27

to make them safe for general use.

01:15:30

And obviously at the same time, in the model level also,

01:15:32

we would have to make some of these changes to incorporate these

01:15:35

changes that happened in hardware.

01:15:36

So it's like a story for both sides.

01:15:39

Is there any data that feeds these studies today from these?

01:15:43

I'm not super aware, to be honest, yeah.

01:15:47

Thank you. Okay, real final questions before we give

01:15:50

them a water break.

01:15:52

Absolutely, this is the last question.

01:15:54

I know, guys, you're super tired.

01:15:55

Last, last question. Yes.

01:15:59

From your experience, guys, training these BLA models and

01:16:03

extracting egocentric data, DATA.

01:16:06

Any recommendation on the quantity or the duration that you require

01:16:11

to achieve a successful train,

01:16:15

to perform a particular task?

01:16:16

For GROOT, or...?

01:16:17

Yeah, with GROOT, exactly.

01:16:19

Exactly, with GROOT.

01:16:21

Like, how much data you need, basically.

01:16:22

400 hours...

01:16:25

Okay, I think Kalyan wants to answer.

01:16:27

Yeah, for N1.7, we use like 40,000 hours.

01:16:29

So 32K real-world, 8K simulated.

01:16:38

So, I think your question was more like, if I'm going to post-chain

01:16:41

the GROOT for a particular task, how much data you need.

01:16:44

It depends on the task, honestly.

01:16:46

If the task is a simple manipulation task, like

01:16:48

100 demonstrations kind of gets you there.

01:16:52

Anything beyond that, like, oh, you want to generalize

01:16:54

it to different things, that is kind of different from that.

01:16:57

You would use Cosmos to augment the data set in different ways for

01:17:01

achieving those generalizations.

01:17:03

But just to mimic the policy on the GROOT side, you just

01:17:06

need like 100 demonstrations.

01:17:07

100 episodes, perfect. Thank you very much.

01:17:09

Thank you. And have a great rest of the GDC.

01:17:12

Thank you.

# How to Build End‑to‑End Physical AI Systems for Humanoid Robots

Akul Santhosh,Solution Architect, Robotics,NVIDIA

Edith Llontop,Technical Marketing Engineer, Robotics,NVIDIA

AI-Generated Summary of this Video

Rate Now

Physical AI brings intelligence to machines, industrial arms, autonomous mobile robots (AMRs), and humanoids, enabling them to perceive, reason, and act in the real world in real time. These robots go beyond scripted behavior to make intelligent decisions on the fly: moving materials, assembling parts, and collaborating safely with people in dynamic environments.

In this session, experts will walk through practical, end-to-end workflows for physical AI development using NVIDIA Isaac and Omniverse technologies. You’ll learn workflows spanning data collection and curation, digital twin creation, model training and scaling, high-fidelity simulation, and deployment to real hardware. We’ll highlight proven toolchains and real-world examples that developers can apply to accelerate the path to smarter, more adaptable robots.

### Learn More About This Topic

#### Overview:

*   >NVIDIA Blog:National Robotics Week — Latest Physical AI Research, Breakthroughs and Resources | April 2026[](https://blogs.nvidia.com/blog/national-robotics-week-2026/)
*   >Product / Solution: 
    *   Lightwheel Accelerates Physical AI Development With NVIDIA Simulation and Foundation Models[](https://www.nvidia.com/en-us/case-studies/lightwheel/)
    *   Humanoid Robots[](https://www.nvidia.com/en-us/use-cases/humanoid-robots/)
    *   Develop Physical AI Applications - NVIDIA Omniverse[](https://www.nvidia.com/en-us/omniverse/)
    *   What Is a World Model?[](https://www.nvidia.com/en-us/glossary/world-models/)

*   >Resource:Physical AI with World Foundation Models - NVIDIA Cosmos[](https://www.nvidia.com/en-us/ai/cosmos/)

#### Developer Resources:

*   >Technical Blog: 
    *   Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T | July 2026[](https://developer.nvidia.com/blog/develop-humanoid-robot-policies-end-to-end-with-nvidia-isaac-gr00t/)
    *   Inside NVIDIA Halos for Robotics: A Full-Stack Functional Safety System for Physical AI | June 2026[](https://developer.nvidia.com/blog/inside-nvidia-halos-for-robotics-a-full-stack-functional-safety-system-for-physical-ai/)

Share

Favorite

Add to list

PDF 

Events & Trainings:GTC San Jose

Date:March 2026

Level:General Interest

Industry:Manufacturing

Topic:Robotics - Humanoid Robots

Language:English

NVIDIA technology:Jetson,AGX,Isaac,Omniverse,OVX,CUDA-X,Blackwell,Cosmos,DGX Cloud,DGX Spark

Region:

Company Information

*   [About Us](https://www.nvidia.com/en-us/about-nvidia/)
*   [Investors](https://investor.nvidia.com/home/default.aspx)
*   [Venture Capital (NVentures)](https://www.nvidia.com/en-us/startups/nventures/)
*   [NVIDIA Foundation](https://www.nvidia.com/en-us/foundation/)
*   [Research](https://www.nvidia.com/en-us/research/)
*   [Corporate Sustainability](https://www.nvidia.com/en-us/sustainability/)
*   [Technologies](https://www.nvidia.com/en-us/technologies/)
*   [Careers](https://www.nvidia.com/en-us/about-nvidia/careers/)

News and Events

*   [Newsroom](https://nvidianews.nvidia.com/)
*   [Company Blog](https://blogs.nvidia.com/)
*   [Technical Blog](https://developer.nvidia.com/blog/)
*   [Webinars](https://www.nvidia.com/en-us/about-nvidia/webinar-portal/)
*   [Stay Informed](https://www.nvidia.com/en-us/preferences/email-signup/)
*   [Events Calendar](https://www.nvidia.com/en-us/events/)
*   [GTC AI Conference](https://www.nvidia.com/gtc/events/)
*   [NVIDIA On-Demand](https://www.nvidia.com/en-us/on-demand/)

Popular Links

*   [Developers](https://developer.nvidia.com/)
*   [Partners](https://www.nvidia.com/en-us/about-nvidia/partners/)
*   [Executive Insights](https://www.nvidia.com/en-us/executive-insights/)
*   [Startups and VCs](https://www.nvidia.com/en-us/startups/)
*   [NVIDIA Connect for ISVs](https://www.nvidia.com/en-us/programs/isv/)
*   [Documentation](https://docs.nvidia.com/)
*   [Technical Training](https://www.nvidia.com/en-us/learn/organizations/)
*   [Professional Services for Data Science](https://www.nvidia.com/en-us/support/enterprise/advisory-services/)

Follow NVIDIA 

[](https://www.facebook.com/NVIDIA "<util:I18n key=\"Follow GeForce on Facebook\" />")[](https://www.instagram.com/nvidia/?hl=en)[](https://www.linkedin.com/company/nvidia/)[](https://twitter.com/nvidia "<util:I18n key=\"Follow GeForce on Twitter\" />")[](https://www.youtube.com/user/nvidia)

[United States](https://www.nvidia.com/en-us/location-selector/)

*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [Your Privacy Choices](https://www.nvidia.com/en-us/about-nvidia/privacy-center/)
*   [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/)
*   [Accessibility](https://www.nvidia.com/en-us/about-nvidia/accessibility/)
*   [Corporate Policies](https://www.nvidia.com/en-us/about-nvidia/company-policies/)
*   [Product Security](https://www.nvidia.com/en-us/product-security/)
*   [Contact](https://www.nvidia.com/en-us/contact/)

Copyright © 2026 NVIDIA Corporation

Select Location

The Americas

*   [Argentina](https://www.nvidia.com/es-la/ "Argentina")
*   [Brasil (Brazil)](https://www.nvidia.com/pt-br/ "Brasil (Brazil)")
*   [Canada](https://www.nvidia.com/en-us/ "Canada")
*   [Chile](https://www.nvidia.com/es-la/ "Chile")
*   [Colombia](https://www.nvidia.com/es-la/ "Colombia")
*   [México (Mexico)](https://www.nvidia.com/es-la/ "México (Mexico)")
*   [Peru](https://www.nvidia.com/es-la/ "Peru")
*   [United States](https://www.nvidia.com/en-us/ "United States")

Europe

*   [België (Belgium)](https://www.nvidia.com/nl-nl/ "België (Belgium)")
*   [Belgique (Belgium)](https://www.nvidia.com/fr-be/ "Belgique (Belgium)")
*   [Česká Republika (Czech Republic)](https://www.nvidia.com/cs-cz/ "Česká Republika (Czech Republic)")
*   [Danmark (Denmark)](https://www.nvidia.com/da-dk/ "Danmark (Denmark)")
*   [Deutschland (Germany)](https://www.nvidia.com/de-de/ "Deutschland (Germany)")
*   [España (Spain)](https://www.nvidia.com/es-es/ "España (Spain)")
*   [France](https://www.nvidia.com/fr-fr/ "France")
*   [Italia (Italy)](https://www.nvidia.com/it-it/ "Italia (Italy)")
*   [Nederland (Netherlands)](https://www.nvidia.com/nl-nl/ "Nederland (Netherlands)")
*   [Norge (Norway)](https://www.nvidia.com/nb-no/ "Norge (Norway)")
*   [Österreich (Austria)](https://www.nvidia.com/de-at/ "Österreich (Austria)")
*   [Polska (Poland)](https://www.nvidia.com/pl-pl/ "Polska (Poland)")
*   [România (Romania)](https://www.nvidia.com/ro-ro/ "România (Romania)")
*   [Suomi (Finland)](https://www.nvidia.com/fi-fi/ "Suomi (Finland)")
*   [Sverige (Sweden)](https://www.nvidia.com/sv-se/ "Sverige (Sweden)")
*   [Türkiye (Turkey)](https://www.nvidia.com/tr-tr/ "Türkiye (Turkey)")
*   [United Kingdom](https://www.nvidia.com/en-gb/ "United Kingdom")
*   [Rest of Europe](https://www.nvidia.com/en-eu/ "Rest of Europe")

Asia

*   [Australia](https://www.nvidia.com/en-au/ "Australia")
*   [中国大陆 (Mainland China)](https://www.nvidia.com/zh-cn/ "中国大陆 (Mainland China)")
*   [India](https://www.nvidia.com/en-in/ "India")
*   [日本 (Japan)](https://www.nvidia.com/ja-jp/ "日本 (Japan)")
*   [대한민국 (South Korea)](https://www.nvidia.com/ko-kr/ "대한민국 (South Korea)")
*   [Singapore](https://www.nvidia.com/en-sg/ "Singapore")
*   [台灣 (Taiwan)](https://www.nvidia.com/zh-tw/ "台灣 (Taiwan)")

Middle East

*   [Middle East](https://www.nvidia.com/en-me/ "Middle East")

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/). You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control (GPC) signal and recorded your rejection of all optional cookies on this site for this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. To opt out of non-cookie personal information "sales" / "sharing" for targeted advertising purposes, please visit the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies. You can manage these settings in the NVIDIA [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies which overrides at least one of your previous settings. You can manage them in the [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

Manage Settings

Reject Optional Accept All

![Image 2: Company Logo](https://cdn.cookielaw.org/logos/10ddf4ca-c072-45d0-b3ac-eead0ed93db0/6e17f6e4-c77b-4a11-9f34-c107c42e4bfc/7981.png)

Cookie Settings

We and our third-party partners (including social media, advertising, and analytics partners) use cookies and other tracking technologies to collect, store, monitor, and process certain information about you when you visit our website. The information collected might relate to you, your preferences, or your device. We use that information to make the site work, analyze performance and traffic on our website, provide a more personalized web experience, and assist in our marketing efforts.

Under certain privacy laws, you have the right to direct us not to "sell" or "share" your personal information for targeted advertising. To opt-out of the "sale" and "sharing" of personal information through cookies, you must opt-out of optional cookies using the toggles below. To opt out of the "sale" and "sharing" of data collected by other means (e.g., online forms) you must also update your data sharing preferences through the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/).

Click on the different category headings below to find out more and change the settings according to your preference. You cannot opt out of Required Cookies as they are deployed to ensure the proper functioning of our website (such as prompting the cookie banner and remembering your settings, etc.). By clicking "Save and Accept" or "Decline All" at the bottom, you consent to the use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) in accordance with your settings and accept our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). For more information about our privacy practices, please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/).

Required Cookies

Always Active

These cookies enable core functionality such as security, network management, and accessibility. These cookies are required for the site to function and cannot be turned off.

Cookies Details

Performance Cookies

- [x] Performance Cookies 

These cookies are used to provide quantitative measures of our website visitors, such as the number of times you visit, time on page, your mouse movements, scrolling, clicks and keystroke activity on the websites; other browsing, search, or product research behavior; and what brought you to our site. These cookies may store a unique ID so that our system will remember you when you return. Information collected with these cookies is used to measure and find ways to improve website performance.

Cookies Details

Personalization Cookies

- [x] Personalization Cookies 

These cookies collect data about how you have interacted with our website to help us improve your web experience, such as which pages you have visited. These cookies may store a unique ID so that our system will remember you when you return. They may be set by us or by third party providers whose services we have added to our pages. These cookies enable us to provide enhanced website functionality and personalization as well as make the marketing messages we send to you more relevant to your interests. If you do not allow these cookies, then some or all of these services may not function properly.

Cookies Details

Advertising Cookies

- [x] Advertising Cookies 

These cookies record your visit to our websites, the pages you have visited and the links you have followed to influence the advertisements that you see on other websites. These cookies and the information they collect may be managed by other companies, including our advertising partners, and may be used to build a profile of your interests and show you relevant advertising on other sites. We and our advertising partners will use this information to make our websites and the advertising displayed on it, more relevant to your interests.

Cookies Details

Cookie List

Clear
*   - [x] checkbox label label 

Apply Cancel

Consent Leg.Interest

- [x] checkbox label label

- [x] checkbox label label

- [x] checkbox label label

Decline All Save and Accept

[![Image 3: Powered by Onetrust](https://cdn.cookielaw.org/logos/static/powered_by_logo.svg)](https://www.onetrust.com/solutions/consent-and-preferences/)

Copy debug info
