---
格式版本: 2
标题: "CUDA: New Features and Beyond S81859 | GTC San Jose 2026"
原文链接: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/"
发布日期: "2026-08-15"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "Published Time: Sat, 15 Aug 2026 07:41:53 GMT"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 4
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-16T01:42:34+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-16T01:40:07+08:00"
入库时间: "2026-08-15T17:42:34.449Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.nvidia.com/gtc/"
匹配关键词:
  - "GPU"
  - "NVLink"
  - "roadmap"
  - "performance"
  - "latency"
  - "bandwidth"
  - "AI"
相关厂家:
  - "NVIDIA"
  - "Meta"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 25
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为CUDA软件特性及Green Contexts技术介绍，未涉及超节点、AI Rack、机柜级系统、供电、散热、互连或量产落地等主题，与项目关注范围无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-16T01:42:46+08:00"
AI主题相关性: 0
AI来源权威性: 15
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 10
采集批次: "2026年8月15日21点37分12秒"
采集批次ID: "20260815-213712-121"
去重键: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859"
---

Title: CUDA: New Features and Beyond S81859 | GTC San Jose 2026

URL Source: https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/

Published Time: Sat, 15 Aug 2026 07:41:53 GMT

Markdown Content:
Visit your regional NVIDIA website for local content, pricing, and where to buy partners specific to your country.

[Continue](https://www.nvidia.com/)

[**GTC Berlin** October 20–22](https://www.nvidia.com/en-eu/gtc/) | [**GTC 2027** March 15–18](https://www.nvidia.com/gtc/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#) 
*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)      
*   [](https://www.nvidia.com/en-us/account/)
*   [Log In](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)[LogOut](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

PLATFORMS

other links

[](https://www.nvidia.com/gtc/)

 Keynote 
*   [Keynote](https://www.nvidia.com/gtc/keynote/)
*   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

 Explore 
*   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
*   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
*   [Speakers](https://www.nvidia.com/gtc/speakers/)
*   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
*   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

[Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)

 More 
*   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
*   [Contact Us](https://www.nvidia.com/gtc/contact/)
*   [FAQ](https://www.nvidia.com/gtc/faq/)
*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*    Keynote 
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*    Explore 
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*    More 
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/# "Menu")

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/# "Menu")

*   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81859/#)
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

[Video 1](blob:https://www.nvidia.com/3cbbbb45-df44-42ce-a8ab-5f1562398e40)

Loading

Play

00:00

Play

Seek 10 seconds backwards

Seek 10 seconds forward

00:00 / 00:00

Mute

Press TAB to open volume control Use the arrows to control the volume

Settings

Play on TV

Turn on Picture in picture

Show Full screen

Transcript Powered by AI

X

X

00:10

Hi everyone.

00:12

Thank you for coming all the way over across the road to come

00:15

and listen to me speak about CUDA.

00:18

I talk about this every year and there's always so much to say.

00:23

And as always I'm standing here representing hundreds

00:25

of people who've done an incredible amount of work and amazing things.

00:29

And so I'm the spokesperson rather than the person behind

00:35

all of these things.

00:37

Typically with these talks I start by sort of presenting how

00:42

I think about, how we think about the landscape and the nature of

00:47

GPU computing and then that helps tie it into where we're going.

00:50

And I talk frequently about parallel programming and

00:54

parallel algorithms.

00:55

Obviously it's what I do.

00:57

This is actually a slide I used at a talk a couple of

00:59

years ago to talk about the way that different forms of

01:03

parallelism get exploited in

01:07

GPU algorithms in deep learning, in scientific computing, and so on.

01:11

But I don't want to talk today about parallel algorithms.

01:15

I'd actually rather talk about how they run.

01:20

Those parallel algorithms have to run on systems that

01:22

are huge, huge like this.

01:24

And so it's not just a matter of how you break your program

01:26

down so that you can divide your data and work on it together.

01:29

It's also how you run this data, how you run these programs, where

01:33

you run these programs and so on.

01:38

A microcosm of that giant data center can

01:40

be thought of as the GPU.

01:42

This is the Blackwell GPU right here.

01:44

It has 160 SMs.

01:46

It has hundreds of thousands of threads.

01:49

This alone can be thought of as a supercomputer.

01:52

I looked it up, it's as big as the largest supercomputer

01:54

in the world from 2004.

01:56

So it's a really big, powerful thing in one chip, and of

01:59

course these data centers are up to 100,000 or more

02:03

of these things these days.

02:05

But what I want to talk about in terms of parallel execution

02:08

and parallel algorithms is something called Symmetric

02:11

Parallelism compared with Asymmetric Parallelism because I

02:13

see a sea change coming in the way that programs run on these systems.

02:19

Symmetric Parallelism I'm defining as when a parallel program fills

02:23

an entire machine and that every single core of that machine, every

02:26

single node, is really running the same thing at the same time.

02:29

Asymmetric is, of course, the converse where you're

02:31

running multiple different things.

02:33

So, in CUDA terms, CUDA has had symmetric

02:36

execution for a long time.

02:37

We have what's called a grid launch in CUDA, where you

02:40

take one program and you run it on all of these SMs, the 160

02:43

SMs in the case of Blackwell Ultra.

02:47

And so you run the same thing everywhere and this is this

02:49

very symmetric mechanism so that if I'm running some sequence of work,

02:55

A followed by B followed by C, then first I run A and everything is

02:58

running A and then B runs and then C runs and it's symmetric because

03:01

everybody's running the same thing.

03:02

It's also sequential because I'm doing all of my A processing

03:05

before my B processing, right?

03:09

CUDA has offered, almost since the beginning, the

03:13

ability to break this up and opportunistically get more done.

03:17

Sometimes the GPU has elements of it which are idle, so we

03:20

have a thing called CUDA Streams, which allows you to say, here

03:23

is lots of different work, you can pick and choose what

03:25

you want to run, and you can create this sequence of operations.

03:30

And several years ago, we took that one step further

03:33

to define a task graph, which encapsulated all of that work in

03:37

a single workflow and says, here is all of the things that are going to

03:40

run and how they're going to run.

03:42

But all of these things do not necessarily fill the whole machine.

03:47

So I can look at this stream, and I can say, well, take a time slice

03:50

through this, and it appears that at this particular time, I might

03:54

be having some A's and some B's and some C's running on my machine.

03:58

But the nature of the way the GPU runs work is that the streams

04:02

merely describe what can be done.

04:04

If A3 jumped in first, it might have filled the whole system

04:07

and you would not be having B2 and C1 running alongside of it, right?

04:12

So even though the parallelism is asymmetric, it's not really

04:16

a guarantee, it's opportunistic.

04:19

And so one of the things I wanna talk about is this move.

04:23

To requiring more of a guarantee.

04:26

And this move is happening because an enormous amount

04:28

of performance and optimization opportunities available if you

04:32

can describe your work in this way.

04:35

So what I have here is a very figurative image of

04:40

of an inference workload.

04:41

Typically you think of this as a production and consumption

04:44

style, you've got a pre-fill where you take in the query, you take in

04:50

the request from the user and you look at your giant amount of data

04:55

and you process it to create the context, to create the thing that

05:00

the decode is going to work with.

05:01

The decode loop iterates rapidly, generating tokens.

05:05

Now, but it's a beautiful description of where the difference

05:12

is between symmetric and asymmetric parallelism and where things

05:14

are moving these days.

05:16

So on the left-hand side, what you have is traditional

05:19

serving, if you like.

05:20

You have a pre-fill operation.

05:23

And then you have the decode operation generating

05:25

the tokens after it.

05:26

Now, the challenge that you have here is if you look at

05:29

that picture on the left-hand side, the pre-fill is very compute-bound.

05:33

It takes an enormous amount of compute work.

05:35

It's doing a lot of matrix operations.

05:37

It's doing matrix times matrix times matrix.

05:39

And so you need massive compute horsepower.

05:42

Whereas on the right-hand side, the token generation,

05:46

the decode component, doesn't need lots of compute horsepower.

05:49

What it needs is lots of ability to iterate fast through memory.

05:52

It needs a lot of memory bandwidth.

05:54

And so the way that you would provision or configure these

05:56

two systems is quite different.

05:58

So the idea on the right-hand side is to say, well, let's

06:00

not try and put this all in one system, sequentially,

06:03

if you remember my ABC from before.

06:05

Let's intentionally run both of these at the same time,

06:08

configured somewhat differently.

06:10

And the benefit can be huge, it can be a factor of 10 or

06:12

even more in terms of speed up to being able to do this,

06:15

because now you've right-sized, if you like, both of the pieces.

06:19

So in terms of how this would run, I've got my Prefill Worker

06:24

running on one set of GPUs, and I've got my Decode Worker

06:28

running on a different set.

06:30

And that allows me this configuration control.

06:32

Now, all of this needs to be overseen by an enormous

06:35

amount of complexity and system just to manage what's running

06:39

where, because these workloads are very unpredictable, right?

06:42

They move around, they change, you don't know what the

06:44

query is going to be.

06:45

There's a lot of different things going on.

06:47

And so, this orchestrator, I've filled in a Dynamo system there,

06:52

which is NVIDIA's orchestrator of this kind of disaggregated system.

06:58

Its job is central control.

07:00

Its job is to manage what runs where and to do it dynamically.

07:04

It's going to spin up these inference engines that are

07:06

going to run Prefill Workers and Decode Workers, and all

07:09

of these things are going to be delicately choreographed together.

07:13

And what's interesting is it's going to depend on the workload

07:17

exactly how this gets orchestrated.

07:19

If I've got a very context-heavy workload, if someone says, here

07:22

is Wikipedia, summarize it for me, that's a massive amount of data.

07:25

I'm going to be very heavy on that Prefill section, on

07:28

the blue section.

07:28

Thank you very much for your attention. But if instead

07:30

I have something where I'm reasoning and I'm generating tokens

07:32

and feeding it back into itself, I'm going to be wanting token

07:35

generation prioritized, right?

07:37

And so what you end up with is this very careful dynamic

07:40

control, very problem-specific, even dynamic within a problem,

07:45

to control that balance between production and consumption.

07:47

The goal is to not starve any component,

07:50

that everything is moving at full tilt all of the time.

07:55

So let's bring this back to the single GPU, because as

07:58

I said, the single GPU itself is a massive, massive system of 160 SMs.

08:02

And I can think about this in the same kind of way, right?

08:08

I can think about how can I take this GPU with teraflops of

08:12

performance and hundreds of SMs and hundreds of thousands of threads.

08:17

You can easily imagine that partitioning that up in a

08:19

similar way, dedicating pieces to specific tasks,

08:24

could have the same kind of benefit that the big system

08:26

desegregation would as well, except in a much smaller context.

08:29

So, while you may not necessarily be able to run an entire AI

08:33

model in the memory of one GPU, there's lots of smaller jobs

08:37

where this kind of partitioning is really interesting.

08:40

Now, the challenge for the GPU has been that even though

08:44

we have a lot of multitasking mechanisms, and they've existed

08:46

for a long time, streams on the left for CUDA, which I

08:48

mentioned earlier, has been around almost from the very beginning.

08:51

On the right-hand side, I have what we call the multi-instance GPU.

08:54

That has hard partitioning.

08:56

But it does not have the dynamicness that is needed

08:59

in these kind of workloads.

09:00

And in the middle we've got the multiprocess service, which has

09:02

done a lot of duty in this regard.

09:05

And it gives me the ability to partition, but it doesn't

09:08

quite give me the dynamicness.

09:10

There's been certain restrictions there which make it difficult

09:13

to get this guaranteed asymmetry that I'm looking for, this

09:16

level of control coupled with this level of partitioning.

09:20

And so what we've done over the last couple of years is we've built

09:23

this thing called Green Contexts.

09:25

And Green Contexts is a mechanism which sits between streams because

09:29

it's within a single process, within a single application,

09:33

allowing dynamic action responding to what's going on.

09:38

And on the other side, the multiprocess service, which

09:40

gives me the partitioning, which gives me the breakdown

09:43

of resources that I need.

09:45

It fits neatly in between with that ability, if you like, to...

09:50

Be a single application, a single controller, like that Dynamo

09:53

controller was, orchestrating my program and dynamically

09:57

assigning resources to what I need, as against having separate

10:01

processes, which is more about isolating different things and

10:04

running them together on the GPU.

10:05

So here's a quick how would you pick between them, but

10:08

I really want to talk about the green context because

10:10

I think this turns on a new ability which really looks to this

10:13

future of guaranteed asymmetry, of guaranteed parallelism

10:18

that I was just talking about.

10:21

So, to create a Green Context is pretty straightforward, even though

10:25

I have a number of steps here.

10:26

You find out how many GPUs are on the machine, 160 in this case.

10:29

You decide how you're going to partition them, and then

10:31

you create a set of descriptors for these different partitions

10:34

so that I can then use these descriptors to assign these

10:37

Green Contexts to them.

10:38

I've got these descriptors here so that you can do nesting

10:40

of context and other such things Things like that.

10:44

Once you have these descriptions, once you've decided on your

10:46

partitioning, you create the Green Context.

10:48

And the context is just a little bounding sandbox in

10:51

which something is going to run.

10:53

You create a CUDA stream from it, and if you're a CUDA developer,

10:56

you'll understand that this stream is the unit to which

11:00

I submit and dispatch work.

11:02

And then, very importantly,

11:05

You just launch kernels into that stream normally, right?

11:07

And this is really the magic in the green context.

11:10

I can manage the execution, in other words, the

11:14

layout, the mapping, the partitioning, separately

11:16

from the individual work, right?

11:20

It's the same as I've got a control managing what's

11:23

running and the things that are running, which are just

11:25

going ahead, doing their thing.

11:26

I can create one of these streams.

11:27

I can launch a...

11:29

A BLAS operation or a Fourier transform or some other

11:31

program, and it will fill half or two-thirds or whatever

11:34

of the GPU, completely oblivious to the fact that I've given

11:38

it a smaller amount of resources.

11:42

So this gives me the ability to do the multiplexing that

11:44

I am looking for.

11:48

All of this comes, because it's a part of CUDA, with

11:51

developer tools which give you deep insight into what's running where

11:55

and how your program is operating.

11:57

Because honestly, if you're doing this kind of parallel

12:01

program control, you've got to be able to inspect what

12:03

it is that you're achieving so you can tune, you can optimize,

12:06

you can see where problems arise.

12:12

This ability to dynamically control what runs where without

12:18

having to impact the work that I'm actually running gives rise

12:22

to a lot of different patterns.

12:23

This is four patterns I draw on a slide.

12:25

You could easily imagine another dozen patterns that

12:26

you could come up with.

12:28

But the key part here is there's all these different things

12:31

you could do with it, because now that I can say what runs

12:33

where and what overlaps with what.

12:35

And it opens the door to things, low latency reservation is a

12:39

common thing, it's an improved form of priority where I've guaranteed

12:45

that resources are available.

12:46

Overlapped execution, nested context, but really the thing

12:49

I want to talk about most is dynamic partition because

12:51

that's the thing that points at this disaggregated workload

12:56

that I've been talking about.

12:58

The ability to control what runs where, for our asymmetric

13:01

parallelism discussion, the ability to control it is key.

13:06

It's the key thing that the green context is doing for me.

13:09

And the way that this works is that these different

13:11

streams that I have,

13:14

Anything put in that stream runs only in the sandbox of

13:18

the green context in which it's been targeted, from which

13:22

the stream has been created.

13:23

And so my stream A is living in box A, my stream B is living

13:27

in box B, and my stream C is living in box C. It seems

13:30

very straightforward and obvious, and I hope it is, because the best

13:34

engineering is invisible like that.

13:37

Another thing you can do with CUDA, I mentioned graphs before,

13:39

you can capture a work of a stream into a single graph.

13:43

And so in the same way that a single graph can span multiple

13:45

streams, if those streams now target individual green

13:48

contexts, I now have a graph that captures work that runs

13:52

in different places, right?

13:53

This is key. I can define a single workflow that spans

13:57

all of my resources and that I can launch from one place.

14:01

Now we're really thinking about this orchestration,

14:03

about this control, the centralized control again.

14:06

This is actually not new for graphs, you've always been

14:10

able to create a graph that spanned multiple GPUs in a system,

14:13

in one node, because those GPUs are connected by NVLink or PCIe, and

14:18

it's funny because we've had graphs for quite a long time, and when

14:21

I mention this people say, really?

14:22

Graphs can run on multiple GPUs?

14:24

And the answer is yes, it's always been able to do that.

14:28

And so if I have a graph like this ABC, I can have A and

14:30

B and C be different GPUs, as well as different green contexts, right?

14:34

The scale doesn't really matter, the graph just describes where

14:36

things want to work.

14:38

Now, of course, if I look a little bit ahead here, that single node

14:44

is almost always part of a cluster.

14:46

It's part of a rack, it's part of a system, right?

14:48

You can easily imagine that if I've got NVLink connected

14:50

GPUs that I can have a graph spanning, then maybe I want a graph

14:54

to just span my NVLink 72 rack.

14:56

Why not? It seems pretty obvious.

14:58

It's not rocket science to imagine that this could happen.

15:00

Now, this is not something you can do today.

15:02

I'm looking ahead to things that we would like to do, because

15:05

I want to tell you about, well, if we would like to do this, what

15:07

are the things which must exist?

15:10

So if I have a graph that can span all of the 72 GPUs

15:16

in one of my NVLink 72 racks, well then, why wouldn't I

15:20

want to span multiple racks?

15:22

Or why wouldn't I want to span an entire data center?

15:24

It seems pretty obvious to me that you'd want to determine

15:27

what runs where on every GPU.

15:29

I want to be able to define a workflow and assign different

15:32

GPUs and resources.

15:33

Think back to my picture earlier of what Dynamo was doing and

15:36

say, here's a bunch of these pre-fills, here's a bunch

15:38

of these decodes, manage them in these different ways, and having

15:41

a vocabulary, a way to say this.

15:43

CUDA is not going to make the decisions that Dynamo

15:45

makes because it doesn't know enough about the system, but giving

15:47

it the tools so it doesn't have to do so much of the legwork itself.

15:51

That's really interesting, but to make this, what seems a

15:54

pretty obvious thing, work requires actually some very sophisticated

15:57

things under the covers, and that's what we're thinking

15:59

about at the system software level.

16:01

There's going to have to be some core multi-node things

16:05

that CUDA is able to do.

16:08

In order to enable this fairly simple and straightforward

16:11

picture of a graph of work spanning all the GPUs that I want.

16:15

Excuse me.

16:17

And the first of these is naming, right?

16:20

This, quick aside by the way, this is the worst picture

16:23

I've ever drawn in PowerPoint.

16:27

It's a 36-node RAIDX4 Dragonfly network, and I asked Nanobanana

16:33

to draw it, and it came out with something comedically ridiculous.

16:36

So I was forced to draw this by hand.

16:38

There's a lot of lines and clicking and things involved here,

16:40

but it is important because...

16:43

When I've got a big system, everybody has to agree on

16:46

what node B3 is, and F2, and if I say run this on D1, everybody's

16:53

got to agree what D1 is.

16:54

If I say move my memory from here to there, we've all got

16:57

to agree on the naming.

16:58

Topology and naming is fundamental.

17:00

It's also extremely hard to do in a consistent way unless

17:03

you're the system software layer, because CUDA is running on

17:05

every single one of these machines.

17:07

And CUDA can sit here and say, yes, we all know who F2 is.

17:10

But that's really difficult to do if you're 10 different applications

17:13

all enumerating it differently.

17:14

So these are the kind of things that need to live

17:16

at the system software layer.

17:18

Now, we already talked about multi-node execution, right?

17:21

Obviously, I want a multi-node CUDA graph to be able to be

17:25

launched from a single location but span GPUs anywhere in my system.

17:29

But it needs a little bit more, right?

17:31

A CUDA graph would need, in this scale of context, to

17:34

be able to make a few additional guarantees, things that we

17:37

really care about, for being able to do our guaranteed asymmetric

17:42

parallelism that I started talking about before, right?

17:44

I want to know not just where things run, but when they run.

17:49

I need to be able to say, these things are going to

17:51

talk to each other, make sure they're running at the same time.

17:54

Or this one over here has to start before that one over there, because

17:57

this is going to set something up that that thing's going to use.

18:00

So in a way, what I've got is a graph which is saying

18:03

what we call a control dependency today is now acquiring a new set of

18:07

dependencies, to say, constraints, which say, make sure that

18:11

A3, B2, and C1 do run together.

18:13

So I do

18:23

The final thing, which is hard to do, and there's, you

18:26

know, we're still very much thinking about these things.

18:28

These are not things that we have in CUDA yet, but I

18:30

wanted to share with you the kind of ways that we're thinking

18:33

about this, is memory management.

18:34

Memory management is obviously a system-level thing.

18:36

Memory management allocation lives in your Linux kernel

18:39

or your Windows kernel.

18:40

It's not something the application takes care of because it is

18:42

a global thing, and it ties into naming and all these other things.

18:46

But so you might have, so the way we think about it, it's going

18:48

to be a little bit more complex than just allocate me a pointer.

18:51

Because you need to think about more than that, right?

18:53

So first you want to say, well, here's a thing, my flower.

18:57

Let's call it a flower, and everybody agrees that

18:58

this is the flower, no matter what node you're on.

19:01

And the way that you break your work down is up to you.

19:03

I don't know how to break, maybe you want squares, maybe

19:06

you want long things, I don't know.

19:07

You're gonna figure out, based on your problem, how

19:09

you break it down.

19:10

So you have to decide how the flower is spanned and

19:13

allocated and arranged.

19:15

And then you have to decide who gets to see what bit.

19:17

This is a really not obvious thing.

19:19

But if I have a machine of 100,000 GPUs, which is a very

19:22

plausible size these days, then if I make a change and

19:26

I have to tell people about it, if I have to tell 100,000 people

19:28

about it, that takes a long time.

19:30

If I even make an allocation and I have to tell 100,000 things

19:32

about it, that takes a long time.

19:34

So instead, what you want to do is constrain and narrow

19:37

the box of who's talking to who, who can see what, and

19:39

that gives you a lot of opportunity to performance tune your work.

19:44

And then finally, the bottom line of all of this is once

19:46

you've done all of these things, you get to point it to your data

19:48

and you just read it and write it.

19:50

In the same way that the green contexts, a kernel running

19:53

inside a green context doesn't know and doesn't care that

19:54

it's in a green context, here we would make the decision that

19:59

the granularity of multi-node CUDA is not one grid spanning many GPUs.

20:04

It's different kernels across the system.

20:07

The kernel itself is local.

20:10

The grid is local, and the graph covers the entire system.

20:14

And that's a key decision point for us, because it means

20:16

that if I look at the picture of what these stacks look like

20:19

these days, a layered multi-node system software stack, where

20:22

everything, in this case, vertical stacks on every single GPU,

20:27

where the orchestrator is managing everything, it looks all pretty

20:29

here, but it's not really pretty.

20:30

What really happens is everybody bypasses everybody else, and

20:33

they're all talking sideways and up and down, and it gets really crazy.

20:36

Because everybody is just trying to hack things together

20:38

right now, and so the idea is that if I can insert a

20:41

layer, a multi-node CUDA driver underneath, which gives you the

20:45

naming and gives you the ability to manage your memory and gives

20:48

you, we're not going to manage the decisions for you, we're going to

20:51

give you the tools to do what you need to do, then the idea is that

20:55

you generate an interoperability here, you make sure that everybody

20:58

is now talking about node F2 being the same node, right, and

21:02

ultimately this fuses down into...

21:05

Just CUDA.

21:06

CUDA becomes multi-node CUDA.

21:07

So this is where we're looking.

21:09

This is CUDA and beyond, and this is the beyond little

21:11

part of the talk, if you like, but I wanted to really let people

21:14

know how we're thinking and think about it yourselves in terms of

21:17

how does this work with the kind of things that you're doing, because

21:20

honestly, nobody runs anything on a single GPU machine anymore.

21:25

Now, it's not like we've done nothing in CUDA in general

21:28

for multi-node and large systems.

21:29

You know, I talked about Dynamo already, which is a product

21:34

that NVIDIA puts out that manages all of this kind of stuff.

21:38

Underneath Dynamo is other things like GPU direct storage.

21:41

Something called CUFILE is the API for that, and it allows

21:44

Dynamo to be very specific and detailed in how it controls where

21:49

the different KV cache locations are stored, and it also allows

21:53

for low latency transfers directly between the GPU and the storage.

21:57

We've also got a checkpointing mechanism.

21:59

This is one of those really interesting things.

22:01

When you build something in engineering and people use

22:04

it in ways you didn't expect, it's always fascinating.

22:06

And here it turned out that checkpointing, there's this

22:08

great talk by Modal, which I've referenced at the bottom.

22:11

I link talks at the bottom of these slides, by the way,

22:13

relating to any of the things that I'm talking about.

22:17

And the idea of using it for elasticity instead of resiliency

22:20

was a really interesting thing that turned up, and it turned out

22:22

to be a very powerful thing to do.

22:26

And finally, of course, developer tools for this.

22:29

We have to be able to investigate and inspect what's going on, and

22:31

so Insight has cloud development tools for managing this kind

22:34

of thing as well.

22:35

There's a whole website for this and talks coming up on

22:39

that later this week as well.

22:41

So I'm gonna switch gears a little bit, but it is related

22:43

and you'll see how it ties in.

22:45

Last year we announced CUDA Tile and in December we launched

22:48

CUDA Tile and CUDA Tile is an extension to the CUDA programming

22:53

model to describe in more detail what you're working with

22:57

in terms of arrays and tensors.

22:59

You're telling the system more about your work.

23:02

And that turns out to be very powerful.

23:04

And I'm going to look back at Dynamo and disaggregated inference

23:07

for a moment because the inference It's benchmarks for something

23:11

like inference X, which sits on it.

23:13

There's a whole layer of things all the way down through this,

23:17

and one of the core pieces of this is called the FlashInfer library.

23:20

This is the library of highly tuned kernels which do the heavy lifting.

23:23

This is the things inside that pre-fill worker and that

23:26

decode worker box.

23:28

And so one of the things we did with the TILE system is

23:30

we took it and we started porting those kernels that

23:35

are used in the benchmark to TILE.

23:39

Because we wanted to move them to Tile because they're

23:42

more portable, it's easier to develop them, but also we wanted

23:45

to make sure for ourselves that the performance capability was there.

23:48

And so what I've got here is a comparison of the pre-CUDA Tile

23:52

versus CUDA Tile implementations.

23:55

And across the board, you know, you're seeing 90 or better percent.

23:59

Now, this is only about half of the kernels.

24:00

We're still working on this.

24:01

We haven't got it all the way there.

24:03

But it's very encouraging to us that here is some really

24:06

interesting and great performance numbers to say that this tool

24:09

could be beneficial to us because we can get high performance

24:12

and all the other things.

24:13

Let me tell you about all the other things, right?

24:14

So what we did was we picked a couple of the kernels out

24:17

of this and we said let's take these kernels because one of the

24:19

things that a key cornerstone of the TAL model is that the code that

24:23

you write is portable across all these different GPU architectures.

24:27

And what we did was we took a couple of those kernels,

24:29

a multi-head attention kernel and a batch matrix multiply kernel,

24:32

And we ran it with no code change on Ampere, Hopper, Blackwell,

24:38

RTX, GeForce, and on the DGX Spark.

24:42

And what happened was there's an autotuner which picks a

24:44

different tile size.

24:45

Blackwell's a bigger machine than, say, Spark, which is the B10.

24:49

And so the Blackwell tile is bigger than the Spark tile.

24:52

But without change, the kernels just moved from GPU to GPU with

24:57

pretty much an 80, 85% and better.

25:00

Performance curve. Now, on the right-hand side at the bottom

25:01

there, we're still working on this.

25:03

We only just released Ampere last week, so we're still tuning some

25:06

of the Ampere compilation things.

25:07

But the idea here is, if I can get kernels that are performed

25:10

enough for that Inference Max benchmark, at the same time,

25:15

I'd write once and run everywhere.

25:16

Now, you take it from here and you start tuning.

25:19

Then you tune for your local specific architecture.

25:20

You start adding hints into your program and other things.

25:22

So you're still doing work to get to that peak 100%.

25:26

But if you're starting at an 80% or a 90% starting point, that's

25:29

a pretty good place to begin.

25:31

And so this encourages us enough that we are going to continue

25:35

internally working with these tools and building, well, building

25:39

out lots of our libraries in Tile.

25:41

And I'll show you more about that in a moment as well, but part

25:43

of the way this works is through this compilation stack where

25:47

Where previously you had just purely a thread-level compiler,

25:51

we now added a tile-level compiler, which can do more, and Cutlass,

25:54

which is the Tensor Core compiler, which lets you get full low-level

25:57

control over the Tensor Cores.

25:58

So I'll show you a bit about how the compiler stack has evolved.

26:02

On the left-hand side is how the compiler stack traditionally

26:04

was, which was kernels talking to thread-level, we call it

26:07

SIMT, which is single instruction, multiple threads, a thread-level

26:10

programming model.

26:12

And on the right-hand side is where we are today.

26:15

It's a bundle of different things that you can target,

26:17

because if you target the right thing, if you use the right

26:18

tool for the job, you can get a lot of benefits, like some of these

26:22

performance and portability things I was telling you about with Tile.

26:24

You can pick whatever thing you want, but you should pick

26:26

the best one, because it makes your life easier or more performant

26:29

or whatever you need to do.

26:31

Under the covers, CUDA has always managed portability

26:34

between architectures through something called PTX.

26:37

But as architecture-specific Tensor Core-like things

26:39

appeared, it became more tricky to keep things portable.

26:43

The tile layer reacquires that portability story, even

26:47

for the complicated Tensor Core architectures of the GPU.

26:50

And what that means is that my library system looks slightly

26:52

different on top of them, where previously everything was

26:55

bound down to a single thread-level representation of my program.

26:59

Now, my library can pick and choose.

27:01

Cutlass has disappeared on the right here because it's

27:03

turned into a compiler.

27:04

But something like FlashInfer, which can pick and choose

27:07

all of the things it talks to.

27:09

And now, this is actually really important, because

27:12

if I can pick and choose what I talk to, we can take some

27:15

of these FlashInfer benchmarks.

27:16

In this case, this is the DeepSeq R1 benchmark for prefill.

27:20

On one of the inference X benchmarks.

27:23

And instead of just porting 100% of the kernels to tile, that's that

27:28

diagonal line bar on the right-hand side which shows you have

27:32

the same performance that you had in tile mode or not in tile mode.

27:36

It turns out that if you can pick and choose, if you're

27:38

sophisticated enough or you want to put the time in, it

27:40

turns out that some of the kernels are much more of a win than others.

27:43

And so you can mix and match these kernels by using this

27:47

different choices of compiler Thank you for watching!

27:50

To get yourself, I mean, in this particular case, a 17%

27:52

improvement on a benchmark by swapping a particular kernel

27:56

from one style to another.

27:58

That's a huge power to put in the hands of your tuning systems or

28:03

your ninja developers, or you can just let the system go and do what

28:08

it will, and you still end up with, in this case, an exact break-even

28:12

performance story, except now you're going to be portable.

28:16

Now, internally, we keep working on this, and one

28:19

of the things we're coming out with imminently in the next

28:22

CUDA release is a C++ CUDA tile.

28:25

We had CU tile in Python, which was CUDA tile in the Python language

28:30

first, that came out in December.

28:32

We're putting out C++ in the next CUDA release.

28:35

And the C++ story looks very similar, intentionally.

28:39

It comes through nvcc, which is our C++ compiler, but we've

28:43

made sure that the

28:45

The C++ sort of idioms, while looking C++-ish, you still have

28:49

the same language and the same look and feel between Python and C++.

28:54

I strongly believe your intuition should just map across no

28:58

matter what language you want.

29:01

And a really interesting thing we did with this was we sat

29:03

down, and don't worry about the big wall of text on the

29:05

left, but I wanted to show you, we implemented an algorithm that does

29:09

not look like the kind of thing you would put in tile mode, right?

29:12

You would, this is a sparse matrix vector multiplier,

29:15

you've got to deal with sparsity, it's implementing something

29:17

called the sliced ELPAC algorithm, and what we did, we wrote a kernel,

29:24

Implementing this algorithm, and we ran it on a test matrix

29:27

suite which Cuspars uses.

29:30

Now sparse matrices are all complicated and there's many

29:32

different categories of them.

29:34

But in a category of these matrices, our one kernel beat

29:38

the state of the art Cuspars library because the compiler

29:41

could bring a lot more to bear.

29:43

We actually looked at this number and we said, wow, that's

29:44

a lot more than we expected.

29:47

All of the things above that yellow line is where our one

29:50

kernel is doing better than And a library of dozens of kernels.

29:55

And we looked in more detail, and we found, well, what is different?

30:00

It turns out that if I take the tile compiler and I express

30:02

my problem in this higher-level array-based model, the compiler

30:07

can do a lot more with it.

30:08

And so I can put one kernel where previously in this specific

30:12

case, doing it in thread Spread level would require 12.

30:17

And so something which is 78 lines of code long can

30:20

perform equally to 1,000 lines in 12 different kernels.

30:25

That is the power of that compilation stack I was telling

30:28

you about, the ability to target a compiler when your problem

30:31

can be mapped onto it, the compiler can do so much more for you.

30:34

Everyone would rather maintain 78 lines of code

30:36

than 1,000 lines of code.

30:38

You'd all like one algorithm that maps, and not only that,

30:41

is portable between architectures.

30:43

Right, so this is actually a really powerful thing that we found.

30:46

And so part of all of these things and figuring out what's

30:49

going on is, of course, analysis.

30:51

Tile is inside inside compute, so you can analyze what's

30:56

going on with your tiles.

31:00

In terms of a timeline, from 13.0, which will run Tilecode, to

31:06

the release in 13.1, 13.2 came out last week with Ampere support, 13.3

31:11

brings in C++ and Hopper support.

31:13

You can all flick through these slides, I'm not going

31:14

to read them all to you.

31:17

But there's also a Connect with the Experts where you

31:19

can come and discuss with us what's on the roadmap and

31:22

what it means to you and how it applies to your applications.

31:26

One of the core pieces, though, whenever we build anything

31:28

in CUDA is this, if you saw Jensen's talk, his keynote,

31:34

he talked about the installed base, the existing pile of things that

31:37

are already out there, the ability for something new to interoperate

31:40

and work with something old.

31:42

And so we built a quick test example where we're calling

31:45

into the FFT library, we're calling into our parallel algorithms

31:50

library, CUDA Cooperative in this case, because it's Python.

31:55

And to make sure that bouncing between these two things incurs

31:59

no real overhead.

32:00

In this case, of this FFT example, the worst case was

32:02

just a few percent.

32:04

Right, so the ability to interface and interoperate with things

32:07

is critical, otherwise your system becomes much more difficult to use.

32:12

We also recognize back in my multi-node theme that communication

32:17

is critical in any application now.

32:20

And so we took a multi-node FFT, Fourier Transform, don't

32:24

worry about the code again, it's just broken into a few

32:26

different sections here in case you enjoy reading code on slides later.

32:29

But really we're just doing a Y then a Z then

32:31

an X Fourier Transform with some communication in between.

32:34

And the communication is the key because we were able to

32:35

embed something called NVSHMEM which is one of our communication

32:38

libraries into the tile program.

32:41

To orchestrate the communication between the

32:43

Fourier transforms that I'm doing.

32:45

And the result was actually, again, a bit of a surprise to us.

32:49

It beat the existing multiprocess FFT by 10, or in the case

32:55

of 64-bit, more than 10%.

32:58

A lot of that came from the ability to interleave the communications

33:01

with the computation at the level that I was just showing you.

33:04

But this ability to make a small piece of code, that

33:09

one piece of code I just showed you, which is fitable on a

33:11

slide although not very readable, to have that produce results.

33:16

In this, like this, has really convinced us that this is

33:19

something that we internally are going to be investing in.

33:21

Our libraries are going to be working on interfaces to

33:25

Tile, programming some of their own stuff up in this Tile

33:28

format, because the flexibility and the performance is there.

33:32

Likewise, our Tensor Core level programming model, which is

33:35

Cutlass, which was the third box on my compiler picture for you, also

33:39

integrating these communications, inline communications inside

33:42

your kernel is critical to being able to manage low latencies and

33:46

control the fine-grained asymmetric parallelism that you're using.

33:50

And so in terms of our communication libraries

33:52

as a whole, all of them are moving towards device

33:56

APIs for communication libraries, as well, of course, as the

33:59

host APIs that they already have.

34:01

This is a slide that I stole from my colleague Matt Nicely's talk.

34:04

He's got a whole bunch of really interesting stuff about

34:06

that as a link to his talk on the bottom there.

34:10

But in terms of libraries themselves, as I said, we,

34:14

NVIDIA, are saying this is really interesting enough that

34:17

we are going to start implementing our own things in Tile.

34:19

We're going to start building Tile interfaces to things.

34:22

The device extension libraries, where you call the linear

34:27

algebra, Fourier transform, compression, all sorts of things,

34:30

inside your kernel, in the same way that communication in your kernel

34:33

is powerful, so too is being able to call libraries in your kernel.

34:37

Now, the libraries are a complicated thing.

34:40

We have CPU libraries, GPU libraries, device libraries.

34:43

One of the things that's really interesting in the move towards

34:46

Python is that the Python compiler lets us pull all

34:49

these things together into a single unit, which we call NVMath Python,

34:53

that does a lot of the dispatching and heavy lifting for you.

34:57

Right, so you would talk to this, if you're a Python

34:59

application, you would talk to NVMathPython and it would

35:02

pull in the correct libraries.

35:03

And it would do more than that, it would be able to JIT

35:06

them and manipulate them in ways that I'll show you in a moment.

35:08

NVMathPython itself is a piece of our overall Python stack,

35:13

CUDA Python, which we are just about to release at 1.0

35:18

version, which means it's...

35:19

Full release.

35:22

We announced it last year, and over the course of this year, we've

35:24

been working on a lot of different pieces of it, and it's now ready to

35:27

really put that 1.0 moniker on it.

35:29

And there's a whole stack of all sorts of different things

35:31

here, and I definitely recommend my colleague Jonathan Dektiar's talk,

35:37

because he will go into much more detail about this kind of thing.

35:40

But the key thing that we've been emphasizing in Python,

35:42

just like in any kind of CUDA, is peak performance.

35:44

And Python is not normally the thing that appears in

35:46

the same sentence as peak performance, right?

35:48

But peak performance means access to, at low overhead,

35:52

access to these C++-isms or these other constructs like

35:57

CUDA graphs and Fusion.

36:00

One of the powers of Python is because it's runtime interpreted,

36:03

I get the ability to do Fusion.

36:06

At the point where my code is called, right?

36:08

And so I can do fusions between different operators

36:12

to make single kernels.

36:13

And in fact, one of the things that the NVMath library does

36:16

is it looks at these calls that you're making, that bottom

36:18

left-hand corner I've got a little number 1, right next to a single

36:21

instruction, which will look at all your types and it will assemble

36:24

the right call at runtime and just in time compile it for you and give

36:30

you a kernel that runs three times faster because it's able to bundle

36:32

it all into one piece together.

36:35

Along with this, again, I keep coming back to developer

36:38

tools because without them, there's nothing you can do.

36:41

And the developer tools, finally we can debug CUDA kernels in Python.

36:45

We can debug number CUDA, which is the single-threaded

36:48

compiler for CUDA.

36:50

In Python, and you can debug it, which is absolutely fantastic news.

36:54

And you can also bring it into Insight, and so you can

36:58

start profiling and examining and inspecting your code for Python.

37:02

That's actually a much bigger deal.

37:03

If you're a Python programmer, you know how much of a big

37:05

deal and how difficult tools are.

37:07

But the ability to, you just can't write parallel code without the

37:10

ability to inspect what's going on.

37:14

Now, I'm going to take a quick look under the covers of the

37:16

compiler stack, because I want to tell you a little bit about how

37:19

some of this magic I've just been telling you about goes on, right?

37:22

And under that stack, if we expand it a little bit, there's

37:24

lots and lots of layers.

37:25

But the layers I particularly want to highlight to you are

37:28

the dark green boxes, which are the things which I mentioned earlier

37:32

are restoring portability for even very complex TensorFlow codes.

37:36

You get portability, so I can run that one code on Ampere,

37:39

Hopper, Blackwell, Spark.

37:41

You saw that graph, right?

37:42

The ability to do that comes from these layers in the

37:46

compiler stack, but the stack itself is more complex.

37:49

We made the stack complex because we understand and we know that

37:52

the things that we build inside CUDA are very often the targets.

37:57

For other compilers, not just for humans to write code.

38:00

So we've got MLIR layers and we have, you know, some people

38:04

target source code and they're welcome to as well.

38:06

We even have a converter so that if you have Triton code,

38:10

you do not have to port your code.

38:11

You can keep your Triton code and plug it into this backend

38:15

with all the composition and all the tools and all the

38:17

other things that are available.

38:18

You can just take your code and make it part of the

38:21

rest of your stack.

38:23

So the story of how these things are built is looking

38:26

at all the different ways people are going to approach it.

38:29

At the same time, we've been based on LLVM, which is sort

38:34

of a compiler layer for translating from one compiler format to

38:38

another, if you like.

38:40

And we've been working to take the LLVM compiler layer work,

38:45

and if you're a compiler person, you know what I'm talking about.

38:48

To put that upstream, so that we are always at the

38:51

top of tree of LLVM.

38:52

Because there's always been this gap, and up till now

38:55

a very big gap, between the LLVM that CUDA supported and

38:58

the LLVM at the top of the tree.

39:00

Now we are taking our LLVM efforts in what we call NVVM,

39:04

which is the NVIDIA extensions to normal LLVM, and we are putting

39:07

that upstream, so from CUDA 13.2,

39:11

I believe it's 13.2 Blackwell,

39:15

Blackwell compiler is going upstream into LLVM 21.

39:22

On a final note on the compiler, there's other optimizations

39:26

to be had as well.

39:26

The compiler is a complicated beast.

39:28

There's thousands of options and things you can tune, and it's

39:32

difficult for a human to tune it.

39:33

In fact, it's so difficult, we built a machine learning

39:36

algorithm to do the tuning for you.

39:38

We call it CompileIQ.

39:39

And the idea is that you can iterate on a particular kernel

39:43

with a particular compiler version, and you can let the

39:46

system go and autotune, if you like, your compiler settings.

39:49

And the result is that you end up with a file called the advanced

39:52

control file, which you can package along with your compilation

39:56

so that you can control the nvcc compiler and the PTX compiler to

40:01

really fine tune what you're doing.

40:03

And what I love about this is this is free performance.

40:07

Meta has been working with us to try this out in their

40:09

kernels and test it.

40:10

These are numbers from Meta on their Trident benchmark

40:14

on the left-hand side and some of the Helion work that they've

40:16

been doing on the right-hand side.

40:18

And the first thing to notice is every single thing is a positive.

40:22

If you spend the time to do this, you get a speedup.

40:25

I mean, a factor of 5% or even 10% speedup isn't a massive

40:28

thing to do just by turning a knob and waiting for a couple

40:31

of hours for it to tune itself.

40:33

This is huge, we're using it for our internal kernels,

40:36

but I also did want to show you some numbers from Meta and thank

40:39

you very much to them for sharing their numbers with us because these

40:43

are real workloads that they're working on in the real world.

40:50

Now, a final thing for me to say here, I've talked about

40:54

tools all over the place.

40:56

Every time I've talked about a feature, I've

40:59

also talked about tools.

41:00

But I think a tool that you cannot avoid in a toolbox

41:03

today, if you do not have some kind of AI coding assistant agent tool,

41:10

then your toolbox is not complete.

41:13

Every single one of us, at least certainly me, and all

41:15

the people I work with, I hope all of you, use these coding agents all

41:19

the time in your development now.

41:21

It is an incredible force multiplier.

41:23

And so, we have invested in Nsight Copilot, which can

41:27

generate CUDA code.

41:29

And it's an assistant as Visual Studio, sorry, it plugs into

41:34

Visual Studio as an extension to Visual Studio.

41:36

There's a whole talk on this, so please go and see the talk

41:39

by some of my DevTools colleagues.

41:42

We've also integrated Nsight Copilot into Nsight Compute

41:45

to help you with your analysis and your profiling and figuring

41:47

out what's going on.

41:48

Because these agents are so powerful, why would you not want

41:53

to have the ability to access them?

41:56

Also on that note, some other colleagues of mine on the

41:58

research side have been working on a way to benchmark and evaluate

42:03

how good is one of these coding agents, how good is an agent.

42:07

And so they've come up with this very interesting structure

42:09

where you can, where they've produced 250 different benchmarks

42:14

culled from an enormous range of different applications

42:17

to test out these coding agents and say what is the, how do you measure

42:20

the quality of an AI, right?

42:21

It's very, very difficult.

42:22

And so this is their...

42:24

Stake in the ground to say that here is a very well thought

42:28

through, detailed, comprehensive suite of benchmarks for the

42:31

AIs to be tested against to see, you know, am I improving the

42:36

quality of my agent or not, right?

42:39

They have a kernel generation competition.

42:41

There's a QR code here.

42:43

You can go to their website and explore the benchmark.

42:47

Again, all of these links, the PDF will be available online, and

42:50

you can just go and follow them.

42:52

And finally, I love that my colleague, one of the most

42:56

common things that is asked of me is can I recommend a book for CUDA?

43:00

Which might be a bit redundant two years from now in the age of AI

43:03

agents, but right now, people are very interested in books for CUDA.

43:06

And my friend, my colleague, my very long time, and a long

43:10

time lover of CUDA, Dr. Wen Mei Wu, is coming out with a fifth

43:14

edition, I think it's just coming out this week, of Programming

43:18

Massively Parallel Processes.

43:20

Every book, of course, is obsolete the day after it's released, so

43:23

today it's current, get it today.

43:25

Because it's the only day... No, that's not fair.

43:28

But not only that, but he's also having a Q&A, and even

43:32

though I hate plugging my own talks, I'm actually moderating

43:34

a panel of all the CUDA old-timers sitting on it, talking crap about

43:39

each other in the old days, and you should definitely come and check

43:41

that out on Wednesday as well.

43:42

That's the bottom link right here, because you have a chance to see

43:45

Wen Mei and just ask us questions about what CUDA's been like

43:49

over the last 20 years, because CUDA 20 years ago was really

43:53

just finally coming together.

43:54

I think the beta was launched in 2006.

43:58

So all the way through this talk, I've given you a list of links.

44:01

There's always a table of this at the end.

44:02

These are talks of my colleagues who provided all this information

44:05

that I've been telling you.

44:06

I give you two slides on everything.

44:08

These guys are gonna go and tell you about it for 45 minutes.

44:12

These are the talks, if you are a CUDA developer, these

44:14

are the ones you wanna be either going to watch or streaming later.

44:18

That connects with the experts to come and talk to us all.

44:20

I'll be at some of those as well.

44:22

And that's the end, so thank you very much.

# CUDA: New Features and Beyond

Stephen Jones (SW),CUDA Architect,NVIDIA

AI-Generated Summary of this Video

Rate Now

The CUDA platform is the foundation of the GPU computing ecosystem. Every application and framework that uses the GPU does so through CUDA's libraries, compilers, runtimes and language—which means CUDA is growing as fast as its ecosystem is evolving. Presented by one of the architects of CUDA, at this engineering-focused talk you will learn about all that's new and what's coming next for both CUDA and GPU computing as a whole.

### Learn More About This Topic

#### Developer Resources:

*   >Documentation Hub: 
    *   Contents — CUDA Features Archive 13.1 documentation[](https://docs.nvidia.com/cuda/cuda-features-archive/contents.html)
    *   Search — CUDA Features Archive 13.1 documentation[](https://docs.nvidia.com/cuda/cuda-features-archive/search.html)
    *   Contents — Release Notes 13.1 documentation[](https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/contents.html)
    *   Contents — Release Notes 13.1 documentation[](https://docs.nvidia.com/cuda/archive/13.1.0/cuda-toolkit-release-notes/contents.html)
    *   CUDA Runtime API :: CUDA Toolkit Documentation[](https://docs.nvidia.com/cuda/cuda-runtime-api/version-mixing-rules.html)
    *   CUDA Runtime API :: CUDA Toolkit Documentation[](https://docs.nvidia.com/cuda/cuda-runtime-api/structcudaKernelNodeParamsV2.html)

*   >Technical Blog: 
    *   CUDA 13.2 Introduces Enhanced CUDA Tile Support and New Python Features | March 2026[](https://developer.nvidia.com/blog/cuda-13-2-introduces-enhanced-cuda-tile-support-and-new-python-features/)
    *   NVIDIA CUDA 13.3 Enhances GPU Development with Tile Programming in C++, Compiler Autotuning, and Python Updates | May 2026[](https://developer.nvidia.com/blog/nvidia-cuda-13-3-enhances-gpu-development-with-tile-programming-in-c-compiler-autotuning-and-python-updates/)

Share

Favorite

Add to list

PDF 

Events & Trainings:GTC San Jose

Date:March 2026

Industry:All Industries

Topic:Developer Tools & Techniques - Programming Languages / Compilers

Level:General Interest

Language:English

NVIDIA technology:CUDA

Region:

Company Information

*   [About Us](https://www.nvidia.com/en-us/about-nvidia/)
*   [Investors](https://investor.nvidia.com/home/default.aspx)
*   [Venture Capital (NVentures)](https://www.nvidia.com/en-us/startups/nventures/)
*   [NVIDIA Foundation](https://www.nvidia.com/en-us/foundation/)
*   [Research](https://www.nvidia.com/en-us/research/)
*   [Corporate Sustainability](https://www.nvidia.com/en-us/sustainability/)
*   [Technologies](https://www.nvidia.com/en-us/technologies/)
*   [Careers](https://www.nvidia.com/en-us/about-nvidia/careers/)

News and Events

*   [Newsroom](https://nvidianews.nvidia.com/)
*   [Company Blog](https://blogs.nvidia.com/)
*   [Technical Blog](https://developer.nvidia.com/blog/)
*   [Webinars](https://www.nvidia.com/en-us/about-nvidia/webinar-portal/)
*   [Stay Informed](https://www.nvidia.com/en-us/preferences/email-signup/)
*   [Events Calendar](https://www.nvidia.com/en-us/events/)
*   [GTC AI Conference](https://www.nvidia.com/gtc/events/)
*   [NVIDIA On-Demand](https://www.nvidia.com/en-us/on-demand/)

Popular Links

*   [Developers](https://developer.nvidia.com/)
*   [Partners](https://www.nvidia.com/en-us/about-nvidia/partners/)
*   [Executive Insights](https://www.nvidia.com/en-us/executive-insights/)
*   [Startups and VCs](https://www.nvidia.com/en-us/startups/)
*   [Documentation](https://docs.nvidia.com/)
*   [Technical Training](https://www.nvidia.com/en-us/learn/organizations/)
*   [Professional Services for Data Science](https://www.nvidia.com/en-us/support/enterprise/advisory-services/)

Follow NVIDIA 

[](https://www.facebook.com/NVIDIA "<util:I18n key=\"Follow GeForce on Facebook\" />")[](https://www.instagram.com/nvidia/?hl=en)[](https://www.linkedin.com/company/nvidia/)[](https://twitter.com/nvidia "<util:I18n key=\"Follow GeForce on Twitter\" />")[](https://www.youtube.com/user/nvidia)

[United States](https://www.nvidia.com/en-us/location-selector/)

*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [Your Privacy Choices](https://www.nvidia.com/en-us/about-nvidia/privacy-center/)
*   [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/)
*   [Accessibility](https://www.nvidia.com/en-us/about-nvidia/accessibility/)
*   [Corporate Policies](https://www.nvidia.com/en-us/about-nvidia/company-policies/)
*   [Product Security](https://www.nvidia.com/en-us/product-security/)
*   [Contact](https://www.nvidia.com/en-us/contact/)

Copyright © 2026 NVIDIA Corporation

Select Location

The Americas

*   [Argentina](https://www.nvidia.com/es-la/ "Argentina")
*   [Brasil (Brazil)](https://www.nvidia.com/pt-br/ "Brasil (Brazil)")
*   [Canada](https://www.nvidia.com/en-us/ "Canada")
*   [Chile](https://www.nvidia.com/es-la/ "Chile")
*   [Colombia](https://www.nvidia.com/es-la/ "Colombia")
*   [México (Mexico)](https://www.nvidia.com/es-la/ "México (Mexico)")
*   [Peru](https://www.nvidia.com/es-la/ "Peru")
*   [United States](https://www.nvidia.com/en-us/ "United States")

Europe

*   [België (Belgium)](https://www.nvidia.com/nl-nl/ "België (Belgium)")
*   [Belgique (Belgium)](https://www.nvidia.com/fr-be/ "Belgique (Belgium)")
*   [Česká Republika (Czech Republic)](https://www.nvidia.com/cs-cz/ "Česká Republika (Czech Republic)")
*   [Danmark (Denmark)](https://www.nvidia.com/da-dk/ "Danmark (Denmark)")
*   [Deutschland (Germany)](https://www.nvidia.com/de-de/ "Deutschland (Germany)")
*   [España (Spain)](https://www.nvidia.com/es-es/ "España (Spain)")
*   [France](https://www.nvidia.com/fr-fr/ "France")
*   [Italia (Italy)](https://www.nvidia.com/it-it/ "Italia (Italy)")
*   [Nederland (Netherlands)](https://www.nvidia.com/nl-nl/ "Nederland (Netherlands)")
*   [Norge (Norway)](https://www.nvidia.com/nb-no/ "Norge (Norway)")
*   [Österreich (Austria)](https://www.nvidia.com/de-at/ "Österreich (Austria)")
*   [Polska (Poland)](https://www.nvidia.com/pl-pl/ "Polska (Poland)")
*   [România (Romania)](https://www.nvidia.com/ro-ro/ "România (Romania)")
*   [Suomi (Finland)](https://www.nvidia.com/fi-fi/ "Suomi (Finland)")
*   [Sverige (Sweden)](https://www.nvidia.com/sv-se/ "Sverige (Sweden)")
*   [Türkiye (Turkey)](https://www.nvidia.com/tr-tr/ "Türkiye (Turkey)")
*   [United Kingdom](https://www.nvidia.com/en-gb/ "United Kingdom")
*   [Rest of Europe](https://www.nvidia.com/en-eu/ "Rest of Europe")

Asia

*   [Australia](https://www.nvidia.com/en-au/ "Australia")
*   [中国大陆 (Mainland China)](https://www.nvidia.com/zh-cn/ "中国大陆 (Mainland China)")
*   [India](https://www.nvidia.com/en-in/ "India")
*   [日本 (Japan)](https://www.nvidia.com/ja-jp/ "日本 (Japan)")
*   [대한민국 (South Korea)](https://www.nvidia.com/ko-kr/ "대한민국 (South Korea)")
*   [Singapore](https://www.nvidia.com/en-sg/ "Singapore")
*   [台灣 (Taiwan)](https://www.nvidia.com/zh-tw/ "台灣 (Taiwan)")

Middle East

*   [Middle East](https://www.nvidia.com/en-me/ "Middle East")

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/). You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control (GPC) signal and recorded your rejection of all optional cookies on this site for this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. To opt out of non-cookie personal information "sales" / "sharing" for targeted advertising purposes, please visit the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies. You can manage these settings in the NVIDIA [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies which overrides at least one of your previous settings. You can manage them in the [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

Manage Settings

Reject Optional Accept All

![Image 2: Company Logo](https://cdn.cookielaw.org/logos/10ddf4ca-c072-45d0-b3ac-eead0ed93db0/6e17f6e4-c77b-4a11-9f34-c107c42e4bfc/7981.png)

Cookie Settings

We and our third-party partners (including social media, advertising, and analytics partners) use cookies and other tracking technologies to collect, store, monitor, and process certain information about you when you visit our website. The information collected might relate to you, your preferences, or your device. We use that information to make the site work, analyze performance and traffic on our website, provide a more personalized web experience, and assist in our marketing efforts.

Under certain privacy laws, you have the right to direct us not to "sell" or "share" your personal information for targeted advertising. To opt-out of the "sale" and "sharing" of personal information through cookies, you must opt-out of optional cookies using the toggles below. To opt out of the "sale" and "sharing" of data collected by other means (e.g., online forms) you must also update your data sharing preferences through the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/).

Click on the different category headings below to find out more and change the settings according to your preference. You cannot opt out of Required Cookies as they are deployed to ensure the proper functioning of our website (such as prompting the cookie banner and remembering your settings, etc.). By clicking "Save and Accept" or "Decline All" at the bottom, you consent to the use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) in accordance with your settings and accept our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). For more information about our privacy practices, please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/).

Required Cookies

Always Active

These cookies enable core functionality such as security, network management, and accessibility. These cookies are required for the site to function and cannot be turned off.

Cookies Details

Performance Cookies

- [x] Performance Cookies 

These cookies are used to provide quantitative measures of our website visitors, such as the number of times you visit, time on page, your mouse movements, scrolling, clicks and keystroke activity on the websites; other browsing, search, or product research behavior; and what brought you to our site. These cookies may store a unique ID so that our system will remember you when you return. Information collected with these cookies is used to measure and find ways to improve website performance.

Cookies Details

Personalization Cookies

- [x] Personalization Cookies 

These cookies collect data about how you have interacted with our website to help us improve your web experience, such as which pages you have visited. These cookies may store a unique ID so that our system will remember you when you return. They may be set by us or by third party providers whose services we have added to our pages. These cookies enable us to provide enhanced website functionality and personalization as well as make the marketing messages we send to you more relevant to your interests. If you do not allow these cookies, then some or all of these services may not function properly.

Cookies Details

Advertising Cookies

- [x] Advertising Cookies 

These cookies record your visit to our websites, the pages you have visited and the links you have followed to influence the advertisements that you see on other websites. These cookies and the information they collect may be managed by other companies, including our advertising partners, and may be used to build a profile of your interests and show you relevant advertising on other sites. We and our advertising partners will use this information to make our websites and the advertising displayed on it, more relevant to your interests.

Cookies Details

Cookie List

Clear
*   - [x] checkbox label label 

Apply Cancel

Consent Leg.Interest

- [x] checkbox label label

- [x] checkbox label label

- [x] checkbox label label

Decline All Save and Accept

[![Image 3: Powered by Onetrust](https://cdn.cookielaw.org/logos/static/powered_by_logo.svg)](https://www.onetrust.com/solutions/consent-and-preferences/)

Copy debug info
