---
格式版本: 2
标题: "Agentic AI Using Open Source Models: Finetuning, Deployment and Chip Design Case Studies S81702 | GTC San Jose 2026"
原文链接: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/"
发布日期: "2026-09-04"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "Published Time: Fri, 04 Sep 2026 01:44:02 GMT"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 4
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-05T21:00:13+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-05T20:59:34+08:00"
入库时间: "2026-09-05T13:01:28.177Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.nvidia.com/gtc/"
匹配关键词:
  - "deployment"
  - "performance"
  - "latency"
  - "throughput"
  - "GPU"
  - "AI"
相关厂家:
  - "NVIDIA"
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 25
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文为GTC 2026技术讲座实录，主线是开源模型选型、微调、agentic AI部署及芯片设计案例，属于模型/应用层，未讨论超节点、AI Rack或机架级系统。虽来源为NVIDIA官方，但无新机架级硬件、标准、量产或部署信息；CodeRabbit与Chip NeMo案例均属单一业务应用，不能形成可复用机架级基础设施机制。缺少机架级主题，命中'应用与模型效率'否决项，总分最高54；按六维实际评分为25，判非优质。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-05T22:14:04+08:00"
AI主题相关性: 2
AI来源权威性: 12
AI新颖性: 2
AI技术细节: 4
AI商业部署信号: 0
AI完整性: 5
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.92"
AI评分知识库SHA256: "60a2554e31d747e68f935f2311b8b8b5d75eff191eaf73eceffba8e80cfd0753"
AI评分知识库检索词: "[\"NVIDIA\",\"Nvidia GTC大会\",\"https://www.nvidia.com/gtc/\",\"RAS\",\"NPU\",\"GPU\",\"Intel\",\"S81702\",\"GTC\",\"URL\",\"GMT\",\"GTCs\"]"
AI评分知识库命中: "[{\"id\":\"runtime-a63fdca649ea6d74b79483ba\",\"title\":\"View On-Demand\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"NVIDIA\",\"https://www.nvidia.com/gtc/\",\"RAS\",\"GPU\",\"GTC\",\"URL\",\"GMT\",\"GTCs\"],\"rank\":-14.513122868385679},{\"id\":\"runtime-47bb908e0d4e4e5fe3587a31\",\"title\":\"View On-Demand\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"NVIDIA\",\"https://www.nvidia.com/gtc/\",\"RAS\",\"GPU\",\"GTC\",\"URL\",\"GMT\",\"GTCs\"],\"rank\":-14.44099850275713},{\"id\":\"runtime-fa6eb5eb4c45c403e767993d\",\"title\":\"View On-Demand\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"NVIDIA\",\"https://www.nvidia.com/gtc/\",\"RAS\",\"GPU\",\"GTC\",\"URL\",\"GMT\",\"GTCs\"],\"rank\":-13.613192626875914},{\"id\":\"july-correct-0034\",\"title\":\"AMD Fires Back at Nvidia with Helios AI System, Epyc CPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"RAS\",\"GPU\",\"Intel\",\"GTC\",\"URL\"],\"rank\":-9.785752714225707},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"NPU\",\"URL\"],\"rank\":-9.731992161721566}]"
采集批次: "2026年9月5日20点55分44秒"
采集批次ID: "20260905-205544-1daa5f4b"
去重键: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702"
---

Title: Agentic AI Using Open Source Models: Finetuning, Deployment and Chip Design Case Studies S81702 | GTC San Jose 2026

URL Source: https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/

Published Time: Fri, 04 Sep 2026 01:44:02 GMT

Markdown Content:
Visit your regional NVIDIA website for local content, pricing, and where to buy partners specific to your country.

[Continue](https://www.nvidia.com/)

Upcoming: Explore other GTCs: [**GTC Berlin** October 20–22](https://www.nvidia.com/en-eu/gtc/) | 

[**GTC Washington, D.C.** November 30-December 3](https://www.nvidia.com/gtc/dc/)

 | [**GTC 2027** March 15–18](https://www.nvidia.com/gtc/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#) 
*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)      
*   [](https://www.nvidia.com/en-us/account/)
*   [Log In](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)[LogOut](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

PLATFORMS

other links

[](https://www.nvidia.com/gtc/)

 Keynote 
*   [Keynote](https://www.nvidia.com/gtc/keynote/)
*   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

 Explore 
*   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
*   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
*   [Speakers](https://www.nvidia.com/gtc/speakers/)
*   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
*   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

[Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)

 More 
*   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
*   [Contact Us](https://www.nvidia.com/gtc/contact/)
*   [FAQ](https://www.nvidia.com/gtc/faq/)
*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*    Keynote 
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*    Explore 
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*    More 
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/# "Menu")

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/# "Menu")

*   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81702/#)
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

Loading

Play

00:00

Play

Seek 10 seconds backwards

Seek 10 seconds forward

00:00 / 00:00

Mute

Press TAB to open volume control Use the arrows to control the volume

Turn on Picture in picture

Show Full screen

Transcript Powered by AI

X

X

00:07

All right, what is up everybody?

00:09

So my name is Chris Alexiuk, I'm a senior product research

00:12

engineer at NVIDIA and I can't see my slides but I assume

00:17

they're up there so that's good.

00:18

What I'm going to talk to you guys today is about the

00:24

open models in the kind of modern agentic AI stack.

00:29

uh so that's a lot of like words but we're gonna we're

00:31

gonna talk through through a lot of them uh and so let's

00:34

just kick it off okay so we've got an agenda here uh we're

00:38

gonna work through in parts i'm gonna bring up some really uh you

00:43

know excellent technical resources for you guys to uh to learn

00:47

from we've got some case studies to walk through uh and we're just

00:51

gonna largely uh Words to the flow.

00:53

Okay, so first thing, the landscape.

00:57

This is, you know,

01:00

Models right now are not just, like, there's one that's the best.

01:03

I think we're all maybe used to the time where there was one model

01:10

that was, like, the best model.

01:11

And that's just the one you would pick and you would use

01:13

it for everything.

01:14

We've entered into a slightly more modern era where there are

01:19

actually loads of models everywhere all the time for everything.

01:23

And this makes it a much more, you know, difficult to navigate

01:28

landscape than it used to be.

01:30

So it's not just like, oh, I use my favorite model and

01:32

that's the only model I need.

01:33

Now we have a model that does, you know, coding tasks really

01:38

well or does tool calling and instruction following really well.

01:41

We have all these models that really can help us and, you know,

01:47

do different tasks in the stack.

01:50

We also have...

01:52

Two big categories of models right now, which is one is open models

01:57

and then one is frontier models.

01:59

So open models, I think hopefully you've heard a lot about today.

02:04

This is things like, you know, NVIDIA, Nemotron 3 family

02:08

of models, Mistral, Kimi, all kinds of incredible folks

02:13

working on open models.

02:16

You know, their weights are available to you.

02:18

They're adjustable.

02:19

You can change them. You can customize these models.

02:23

You can deploy them in very specific configurations, which

02:26

we're going to talk about a little bit later on in this session.

02:29

But the idea is that these are models that you have full

02:33

access to and control over.

02:37

Now there's a separate category, which is frontier models.

02:41

These models are not so open.

02:43

However, they're usually quite, let's say the word powerful,

02:48

strong, capable, right?

02:51

The idea is that these are models that are operating like at

02:55

There are a number of ways to deploy these companies

03:00

on-prem, and they're typically held behind an endpoint that

03:04

you're going to reach them through.

03:05

Though, obviously, of course, you can work with any number

03:09

of these companies to deploy them on-prem through a number of means.

03:14

The idea is that they are very large.

03:18

Locked down, but, you know, strong models, very capable models.

03:23

And the idea here is that this is a really important

03:29

decision that we have to make in this landscape right now, right?

03:32

Which is, when do I choose open models versus when do

03:36

I choose closed models versus when do I choose something

03:40

just a little bit different?

03:43

There's a simple case for open models, which is that

03:46

if your model meets any number of these criteria, you're basically

03:51

going to want to include somewhere in your stack an open model for

03:56

the easiest time to flight, right?

03:59

You know, there are reasons that you might just need to

04:03

have the weights.

04:05

in your environment.

04:07

There are reasons that you might need to have transparency

04:10

of data lineage to use the model.

04:12

There are a lot of decision points that exist where you

04:15

just have to know exactly what the model is, what it's

04:20

running on, how it's running.

04:21

And in that case, our decision is not very difficult, right?

04:26

We could just kind of say there's only one category

04:29

of models that actually leverages that, so that's what we have to do.

04:34

Every other case is much more difficult, okay?

04:37

And so there's a bunch of trade-offs that we're

04:39

gonna talk about.

04:40

I'm gonna set the table here, right?

04:42

So my job today is to introduce the problem, and then we're

04:46

gonna have some of my colleagues and experts come up.

04:51

We're gonna talk to you about how they approach this problem.

04:54

So for here, we have the table being set with

04:58

cost at scale, right?

05:00

So this is the idea that if we need to deploy many of these

05:04

models or have many users hitting our models, especially for some

05:09

tasks, which are maybe, you know,

05:12

A little bit easier, right?

05:14

They don't require PhD-level intelligence.

05:17

Things like that, right? We can actually deploy open

05:21

models far and wide.

05:23

We can deploy them on the edge.

05:24

We can do all kinds of fun things with open models we

05:27

don't necessarily have access to with the frontier models.

05:29

Latency control, this is a big one.

05:32

Sometimes we just need the model to be super fast.

05:35

We need it to be super fast at this time exactly.

05:39

With some of the frontier models, we have access to

05:42

things like, you know, we can do like super speed mode now, right?

05:45

With Claude and OpenAI and things like this.

05:48

But with our open models, we have direct control of

05:53

that tuning lever.

05:54

We can throw as many GPUs at the problem as we need.

05:57

We can deploy it in exactly the configuration that

06:00

we need for exactly our customers or our use case.

06:03

So we have that, we have our hand on the control knob there.

06:06

Customization, this is a big one, right?

06:08

My problem is not, your problem is not the person sitting

06:11

next to you's problem.

06:12

Customization allows us to change how the model

06:15

behaves up to and including a fundamental level, right?

06:19

So it's not just like a...

06:21

We can change it a little bit.

06:23

We can have it change it through prompts or whatever.

06:27

We actually get the ability to deeply change the model.

06:30

If we need it to be very good at JSON and JSON only, we can

06:34

make the model do that, and that's one of the benefits of open models.

06:38

Thank you to the incredible tech team.

06:40

We've got the screen here now.

06:41

Okay, so the next one is capability ceiling.

06:44

This is a big one.

06:45

Sometimes you need the model to be able to do PhD level mathematics.

06:48

Sometimes you need it to be able to reason through complex

06:50

problems and decompose them into smaller steps.

06:54

And that is just like your use case only works if you're

06:57

able to do that, right?

06:58

And open models are not yet caught exactly up to the frontier, right?

07:03

Which means that we do have to make a decision about capabilities

07:07

when we're thinking about

07:08

The trade-offs between the two families here.

07:12

Operational burden, this is a big one.

07:14

If you wanna move fast and break stuff, there's a lot more

07:18

stuff to break if you are deploying the models yourself, right?

07:21

So you have to own it, you have to be the owner of it.

07:25

And that means either more engineers or you're asking

07:31

somebody outside to help you out.

07:34

Versus frontier models, I mean, we hit the API, it's stable.

07:37

That's the way it goes.

07:38

Usually, at least.

07:40

And then of course, privacy and data sharing.

07:43

You own all the data in and out if you have the model on-prem, right?

07:47

You own everything about the model when it comes to privacy

07:50

and data sharing, frontier models.

07:52

There are lots of excellent compliance paths and everything

07:55

like that, but you wind up in a situation where you have

07:59

the ultimate control if you deploy the model.

08:01

Okay, so what is the answer?

08:04

Well, the answer is, of course, many of you probably

08:05

already knew this.

08:07

We should probably use more than one model, right?

08:09

Like we probably shouldn't limit ourselves to just one of these

08:12

very powerful, intelligent systems.

08:14

We should use a system of very powerful, intelligent systems.

08:18

And the idea here is that I think the days of like, my

08:22

agent runs on, you know, whoever's frontier model or whatever single

08:26

model are nearing an end, right?

08:29

The idea is now,

08:30

We should use models the way that they're intended to be, which

08:35

is for the thing they are best at.

08:37

If we need a certain part of our stack to be very fast, and we don't

08:41

need it to be hyper-intelligent, we can slot open models into that part

08:46

of the stack instead of having the frontier model do the whole thing.

08:51

So if we have kind of this,

08:54

You know, we need to be secure, let's say, right?

08:57

We need to be safe with our data.

08:58

We can have a model that we control be part of that stack

09:03

for that component of work, right?

09:04

And this is the idea that, you know, we're moving away

09:07

from kind of one big monolith model powering everything

09:10

to a constellation of models.

09:13

And then, of course, you can read the slides, so we don't

09:15

gotta do that for you.

09:18

So now how do we pick the right one, right?

09:20

So like, this is the obvious next question, which is we

09:23

have so many models in the world, just the Neotron 3

09:27

family alone, there is a bunch of models to choose from, right?

09:30

So the idea is how do we pick the best one?

09:32

It really comes down to a lot of different...

09:37

You know, decisions that you have to make.

09:40

We're gonna hear a case study on this in just a moment, but just

09:43

to lay them out for you, right?

09:44

We need to know how strong it is at task A. We need to

09:47

know how much it's gonna cost for us to own that over time.

09:50

We need to know how much we can customize it, what we're

09:53

allowed to do with it.

09:55

You have to start getting into licenses specifically, right?

09:58

You have to care about, what am I allowed to do with this model?

10:02

How many customers am I allowed to serve with it?

10:04

All of those things become

10:06

Critical to understand when you're trying to pick a model.

10:10

And then, of course, latency, right?

10:12

There's a big difference between Nemotron 3 Super BF16 and

10:16

Nemotron 3 Super NVFP4, right?

10:19

There's not much of an accuracy degradation.

10:21

In fact, it's like 99 point X percent, right?

10:24

But there is a big difference in performance as in throughput.

10:28

And those are decisions you have to understand in order

10:31

to really choose the right model.

10:34

And then we have this sizing chart which is we have many different

10:39

cases of models and we're going to talk about this when we talk about

10:42

deployment but these are relatively correlated with ease of deployment

10:46

right so it's not hard at all to deploy a nano-sized model.

10:52

Right, you can you can do that on a number of different resources

10:55

and it's going to be just fine.

10:57

But as you move up through those capabilities, which

11:00

means our models are getting better and better at whatever

11:03

tasks we throw at them, we're going to find that it's more

11:07

and more difficult to actually get these things running in the wild.

11:11

And so there's another axis we have to think about when it

11:13

comes to picking our model, which is do we want to take on the burden

11:17

of setting up a ultra-class model, which is like 400 billion parameter

11:22

plus model for our use case.

11:25

And that ultimately is going to be a decision that you

11:28

guys have to make.

11:30

And a company that made that decision is going to tell you about

11:33

how they did it, is CodeRabbit.

11:35

And so I'm going to ask that David Loker, VP of AI, come

11:39

up and explain what they did.

11:43

Thank you very much.

11:46

We got some feedback there.

11:48

All right. Yeah.

11:49

So hi, everybody.

11:51

As I said, David from CodeRabbit.

11:53

I spoke a little bit earlier.

11:54

For those of you who don't know, we are an AI code review company.

11:57

And when we think about model selection, we don't think about

12:03

we're going to pick one model and it's going to take everything on.

12:06

Because our workflow is varied, there's a lot of

12:08

different pieces to it.

12:09

There are varying complexities, and so there's a lot of things

12:12

that go into that.

12:13

So I wanted to talk about how you choose the right model

12:17

for the right use case.

12:20

So obviously I can have lots of different options.

12:22

They can be frontier models.

12:23

They can be open models.

12:25

I have to make a choice and I want to make the best choice

12:28

that I can with the information I have available to me.

12:30

But at the same time, how?

12:33

How do I go about setting up a system so I can actually

12:35

make that choice?

12:38

So in this case,

12:41

This isn't about ideology, right?

12:43

So this isn't about open versus frontier necessarily.

12:47

It's about engineering discipline, right?

12:50

How I approach problem solving.

12:52

So I'm gonna share how we evaluate models, how we evaluate our prompt

12:56

changes even, how we evaluate our context engineering system,

12:59

and anything that might go through into these models themselves.

13:03

So, AI code generation obviously is accelerating, but risk

13:06

is as well, and that's where CodeRabbit comes in, right?

13:09

So we have AI coding speed up, we have PR volume up,

13:13

we have cross-reflow coupling up, and obviously we need

13:16

some sort of quality gate.

13:18

We have lots of problems, right?

13:21

We're not trying to replace developers here, we're

13:22

trying to keep systems as reliable as possible

13:25

as this throughput explodes.

13:29

So our main thesis is you test, you don't assume, right?

13:35

Whether you're doing Frontier versus Frontier, Open versus

13:37

Frontier, anything, whether you're changing your thinking budget

13:40

token counts, whether you are changing your prompts, everything,

13:44

every change is a hypothesis until it's proven to be useful.

13:48

And if you approach it from that way, you can start to

13:49

build the systemic thinking of how I approach this problem.

13:55

Benchmarks is not the same as your product, right?

13:59

If it doesn't move your own key performance

14:03

indicators, it's just noise.

14:04

Public benchmarks may not actually apply to you in any case.

14:08

It might be an indicator, it might give you directionality

14:10

when it comes to what things or models you might be interested in,

14:14

but ultimately you need to test it.

14:17

Right, so your product has its own thing.

14:18

So for us, for example, it's actual repositories, actual PRs, how

14:24

people interact with the product, the latency that might come

14:26

into it, the trust that we have to earn from the customer, right?

14:32

And reliability of our product, that has to perform the same

14:35

over time and make sure you're getting the same value.

14:37

So you have to measure your own outcomes.

14:41

What matters to your business?

14:43

What matters to the use case that you're trying to optimize for?

14:48

And so we go through this

14:51

phased approach to testing when it comes to any of these

14:55

changes that I mentioned, right?

14:57

We have an entire offline system, which is going to catch obvious

15:01

breakages, it's going to help us compare candidates quickly, it's

15:04

going to let us iterate rapidly.

15:07

And when we find some directional improvements, when we find

15:10

things that seem to be moving the needle the right way for

15:12

what we're trying to accomplish, we're then going to move into this

15:16

sort of shadow realm, so to speak.

15:19

We're going to try and prove that the whole pipeline

15:21

behaves as expected, the prompting, the context,

15:24

the routing to the different models that we might choose.

15:28

And once we've then proven that out, then we're going to move

15:31

online and we're going to watch our KPIs and make sure that what we see

15:36

online matches what we see offline.

15:38

And so that's some of the things I'm going to talk about

15:40

in a bit and how you set that up.

15:43

So these are the things that we measure.

15:46

So it's a lot, but these are the things that we have found

15:50

to indicate when they go in the right direction and the other ones

15:54

don't go in the wrong direction.

15:56

That we see positive effects on our user base in the way

15:59

that they enjoy the product and how useful it is.

16:03

So for us, we're catching bugs, right?

16:06

So we have a recall rate, obviously.

16:08

How many bugs did we find?

16:10

I don't know that in an online space because I don't know

16:12

how many bugs are in your software that you're submitting.

16:15

But in an offline sense, I can get a code base and I can

16:18

purposely inject bugs into it or I can find open source repositories

16:23

where someone has fixed a bug and I could find where that

16:25

bug came from and I can test that.

16:28

So in this case, I know the bugs I'm trying to find.

16:31

I have precision of how many comments did it take for me

16:35

to actually find those bugs.

16:37

If that number of comments went way up and I didn't catch any

16:39

more bugs, then I can expect when I move that online, I'm going to

16:42

have a higher noise in my system.

16:45

If I see the number of comments going down, but I'm catching

16:47

the same number of bugs, I can expect my signal-to-noise ratio

16:51

to go up and the acceptance rate of comments online We'll go up.

16:56

But crafting that offline experience is actually

16:59

extremely critical.

17:00

You have to make sure the directionality works.

17:02

So when you're testing your benchmarks, you're building

17:04

your offline evaluation system, and you're picking those metrics,

17:08

really pay attention to all the different ones that are there

17:10

and how they end up playing out online in the real world scenarios,

17:14

because there's a lot more variance in that data than what's going to

17:16

be available in your offline set.

17:18

And we've seen that a lot in benchmarks that aren't

17:20

representative of real world.

17:24

So we have other things too like verbosity, usefulness,

17:27

severity mix, a lot of things that we're watching just to

17:30

make sure that when we choose between different models,

17:34

we're aware of the repercussions or the impact that that would

17:37

have on our system as a whole.

17:41

In online, we can actually do things like sentiment.

17:43

Did a user like it?

17:45

Didn't they? Did they have a really strongly worded comment

17:48

in response to our comment?

17:51

That happens. And whether we can detect hallucinations.

17:54

So we'll say something, the user will say, that is absolutely

17:57

incorrect, and then give us a reason, and we'll know,

17:59

OK, we had a model hallucination.

18:01

And it used to be that libraries was one of the number one

18:04

reasons, and we were able to find that because of the

18:06

collection of this information.

18:08

And then make adjustments and find that that we took that

18:10

information, brought it into our offline and we were able to make

18:13

adjustments to then fix that issue.

18:18

Prompt physics, yeah.

18:22

So different models obviously interpret the same instructions

18:25

very differently.

18:26

If you've ever played trying to take your prompt and switch between

18:29

open AIs, models, anthropics models, and emotrons, or any

18:31

of these models, you'll notice that things behave wildly differently.

18:35

The way that they were trained, the end result of this sort

18:39

of final stages will dictate a lot of the ways that they

18:43

react to certain types of language.

18:45

If they're really tuned towards tools and you use the word

18:48

tools in your prompt, you can expect the thing to take

18:51

a certain turn when it comes to the way that it likes to use tools.

18:55

And so they've been tuned in certain ways,

18:57

so you have to adjust.

19:00

So both in the structure, the verbosity, the calibration,

19:04

what these things flag versus ignores in our particular case.

19:08

So this whole thing is a life cycle.

19:10

Every single time you want to change a model or change a prompt,

19:13

there's a whole process that you have to go through in order

19:15

to make sure what you're doing actually has the impact you want.

19:22

So you kind of have to treat your prompts like software

19:24

modules is the way that I would like to think about it.

19:27

Right, the core might stay stable, fairly stable.

19:30

And then you have these little adapters that work for different

19:32

models and you can test them out.

19:34

So one model, you need to tell it to be much more quiet, needs

19:38

to prune out certain ways that it likes to think about things.

19:41

And so we have very different verbiage that we might

19:44

use for Anthropic versus Nematron versus OpenAI

19:47

models, depending on the use case.

19:50

So this is the thing that makes routing feasible, right?

19:52

I have my core prompt, it's stable, and I have a few other

19:55

ones that are specific.

19:56

Now when I want to route, I have the ability to tack

19:59

those on as it fits for any of those particular models, and I've

20:02

measured them so I know the impact.

20:09

So, bringing a model on safely.

20:13

Firstly, the first step that we go through when we're thinking

20:16

about adding any model, and we thought about this when

20:18

we were trying to look at both Nano and Super, is what do we think

20:24

this model is going to be good at?

20:26

What's the size of the model?

20:29

What are the sort of benchmarks that we can see directionally

20:31

where it sits in the space?

20:33

So we can see maybe how they thought about training that model.

20:36

What did they try to specialize it in?

20:38

So we get some insight into that.

20:40

If they have the prompts or various other data that they

20:44

have available, that's very useful information for us to understand a

20:47

little bit about the model itself.

20:50

So that comes into it, right?

20:51

And then we go through our evaluation phase.

20:53

Okay, now that we know maybe where we should try it out,

20:55

because we don't want to try it out on everything.

20:57

We don't try out the nano model, for example, on our

21:00

heavy reasoning tasks.

21:02

That wouldn't make a lot of sense.

21:03

And so we kind of try and look at our workflow, the different areas

21:07

of where models are getting used.

21:09

And so in the case of nano and super, for example, we're

21:12

looking at things like our context engineering workflow.

21:16

Everything that goes through there, it's not a heavy reasoning

21:18

task, but the idea is that I need to distill information

21:21

from like an MCP document or distill information from a file.

21:25

I might be looking for something very specific.

21:28

And in those cases, these models tend to perform very well.

21:32

But how do I tell that?

21:34

So that's the evaluation part.

21:35

That's where I go through my offline harness.

21:37

I do multiple iterations to try and figure out how does that

21:40

impact all the downstream tasks?

21:43

If I evaluate the outcome of a model of information

21:47

distillation, and I run it up against a bigger model, or one that

21:50

I know to be my current standard, does it find the same pieces?

21:53

Does it find more?

21:55

We noticed the super actually had an uptick over some frontier

21:58

models in comparison at finding and extracting information.

22:03

So you can learn a lot through that offline evaluation.

22:07

You go through adaptation, so I find maybe some holes.

22:10

And I think that they're not huge chasms of gaps that can't be filled

22:15

by the model, but they seem to be just things that are slightly off.

22:18

And so I look at some of the cases that might be astray and look

22:21

at the thinking tokens or various other things that might be there.

22:25

I evaluate them and I adapt my prompts.

22:27

That's the adapter part of it that I was talking about before.

22:29

My core remains the same and I write an adapter.

22:32

And that might fix some of those holes and get me

22:34

a little bit further.

22:36

I do that for a while and then I go through my shadow

22:38

and my rollout phase.

22:40

And then finally we're in some steady state,

22:42

hopefully at some point.

22:43

Whether that happens or not depends on whether they release

22:46

yet another model in a week or so.

22:50

But context engineering, that's our hidden lever, right?

22:55

Most of our failures happen before reasoning, so missing

23:00

evidence, the wrong evidence.

23:02

So I bring in incorrect information.

23:05

I have too much evidence, so the model's unable to pay

23:08

attention to all of it.

23:10

And if the context is wrong, it doesn't really matter

23:14

which model you use.

23:15

It will most likely fail.

23:17

And so that's part of my previous talk, which is

23:20

that context engineering is just about everything

23:22

when it comes to a task as complicated as doing code review.

23:33

So structuring your evidence reduces hallucinations.

23:35

So, for example, we have a lot of information, and this is

23:39

a lot of information in this slide.

23:40

There's a lot of pieces that come into play when it...

23:44

When you're trying to figure out how to review code, and

23:46

even as a human being, I'm trying to review a very large

23:50

PR, and I might just go ask the person, what were you thinking?

23:54

What is this supposed to do?

23:56

Even if I'm familiar with the codebase, I look back

23:59

at something I wrote six months ago, I might forget what exactly

24:02

I was doing at the time.

24:04

So there's a lot of information gathering in order for me

24:07

to do a good job, and the same holds true for any model.

24:13

And so that's the heart of CodeRabbit is that context

24:16

gathering, information gathering, synthesizing that information

24:19

and ultimately optimizing the context so that we don't

24:22

put in too much or put in, yes, maybe factually correct,

24:25

but not necessary information into the context window to have the

24:29

sort of context rot, if you will.

24:32

And so Nematron actually can sit right in that middle,

24:34

and almost 80%, up to 90% of all of our LLM calls will

24:39

be to that central system to help us build the context window.

24:43

So to make sure that information is as necessary and concise

24:47

as possible with all the details that we might need in order

24:50

for the reasoning model, which is maybe Cloud or GPT, to

24:54

actually do the really heavy lifting and figure out what

24:57

comments are necessary, what bugs are there, and things like that.

25:01

So that's a lot of work that's going through that model.

25:04

And the only way we were able to make that choice is through

25:06

our very, very hard-fought evaluation system that's taken

25:12

just over eight months to build to this current state.

25:18

So, I mean, this is something I just touched on to a

25:20

certain degree, but token bloat kills quality.

25:24

I mean, I've seen that again and again.

25:25

And as I've done experiments before, we just put in as

25:30

much information as possible.

25:31

And basically you get confusion, the model just

25:34

loses the plot, right?

25:36

It's suddenly unaware of what it's supposed to do.

25:38

It latches onto something at the end and it just goes

25:40

with it, ignoring sort of instructions in the middle.

25:44

You have clash, you have conflicting facts, so if

25:45

I pull in too much I might end up with something that's

25:48

not congruent with each other.

25:50

And I just might bring in bad sources of information.

25:53

So there's a lot of problems in there.

25:55

So we have to dedupe, we have to make sure we're

25:57

doing things correctly.

25:59

We have to, again, make this into an optimization problem,

26:02

as opposed to just throwing as much at it as possible.

26:08

So one of the interesting things I wanted to bring up

26:09

is the context that comes in for us that's outside the diff, right?

26:13

These are some of the things that Nematron helps with.

26:16

So we're pulling information from the repository and now

26:19

we have multi-repo information that we can gather.

26:22

So front-end might use back-end code.

26:24

How do I know which information is necessary?

26:27

How do I extract that information?

26:29

All these questions and all these iterations and all this

26:31

tool use, gathering information through our knowledge-based system,

26:35

which has past PRs indexed, which has learnings that the user talks

26:40

to Corev and tells us things about how they like to do code review.

26:45

Or even path instructions where up front you know these

26:48

are my coding guidelines, this is exactly how I want variables

26:51

to be named and so on and so forth.

26:53

All that information needs to be gathered up only

26:56

in the right space.

26:57

I don't need to pull in everything if it doesn't apply to this PR.

27:02

This is a small example you can see where someone, let's

27:06

say, has made a change to process refund and that's the diff.

27:09

They added an exception and if you don't know to look

27:12

outside the diff,

27:14

And do some analysis outside, you wouldn't notice the handle order

27:17

doesn't handle that exception.

27:18

And so when that happens, you're going to have an issue.

27:21

And so this is where you need to look outside the diff, you

27:23

need to explore where the contracts are, both input and output.

27:27

In this case, the bug was one call away.

27:34

So evidence-based verification is another area where actually

27:37

the supermodel can be used, right?

27:39

So again, this comes back to

27:43

the idea that we can verify things that are coming through.

27:48

So we analyze things that we verify and then we decide.

27:52

So tool-based checks reduce hallucination.

27:54

Verification loops are going to matter more and more.

27:57

And frontier models are still winning, though, on hard,

28:00

ambiguous checks.

28:08

All right, so how do we route, right?

28:12

So we do a few things.

28:13

So we look at things like token usage, the latency tolerance,

28:19

in our workflow, is this a spot where high throughput or

28:23

low latency is extremely important?

28:25

If it's something that happens a lot and happens in a loop,

28:28

chances are I'm not gonna wanna use a large frontier model because that

28:31

will blow up my total review time.

28:33

In a lot of cases, it's just not necessary.

28:35

You can try it, you put it in there, and the amount of

28:38

benefit you see to the end user is very minimal.

28:40

That goes back to, again, where our evaluation system

28:43

is very, very useful.

28:46

What's the cost profile?

28:47

So we don't pass token costs onto our customers.

28:50

So again, our incentive is to optimize for the quality

28:53

using the model that's appropriate.

28:55

And so is it actually providing a real benefit to a measurable

29:00

degree that can warrant the additional cost?

29:05

And then for a problem in general, what's the complexity?

29:06

What's the ambiguity level?

29:08

Is it something that I think that a frontier model needs to take on?

29:12

Or is it something that I can use something smaller?

29:15

And so all this sort of shape, again, my decisions as I'm thinking

29:18

about which model to route to.

29:22

But everything always grounded in that empirical metrics that

29:27

I can actually follow and rely on.

29:36

So, as I was saying, certain high-volume context subtasks

29:39

perform extremely well on Nematron models.

29:44

Right now, Nematron is supported in our self-hosted deployments,

29:48

and Frontier models remain best in class for deeper reasoning,

29:53

ambiguous edge cases, as well as our verification loop.

29:57

So after things go through, in order to verify that those

30:01

comments are grounded in truth, double checking everything,

30:04

making sure that there's evidence behind stuff, that's still

30:07

on the Frontier model level.

30:11

Although I'm interested in trying out Ultra when it comes

30:14

out to see how it performs.

30:18

A lot of my time actually is spent measuring models,

30:22

evaluating models, trying to figure out whether the extra

30:26

token uses for certain contexts is going to be valuable or not.

30:28

And so this is a process I'm used to.

30:30

And ultimately, if there's any investment that you can make

30:35

in your company, if you're They're going to be adopting any sort

30:38

of models in building an extremely

30:42

Valid eval framework and harness that you can look at those

30:46

metrics and map them to some sort of movement in your online metrics.

30:50

So directionality.

30:52

It doesn't have to map one-to-one.

30:53

I don't need my recall rate to be exactly the same as

30:56

offline and online.

30:57

What I need is that when this one goes up, the other one goes up.

31:00

I need that to remain constant.

31:01

As long as I can keep that in play, I can really rapidly iterate.

31:05

I can make really, really fast progress.

31:08

And that's what you want.

31:15

So open models are compounding, I mean, in our eval super,

31:19

you know, it improves the same context engineering tasks

31:22

where nano was strong.

31:23

So nano also did a good job.

31:25

Super is even better.

31:26

It's really great at instruction following.

31:30

And it also is really good at tool use.

31:32

It follows instructions extremely well.

31:33

Notice it very, it adheres to them more so than, than other models.

31:37

And so again, we had to write an adapter to, to help it

31:40

to get back to where we were before and change things up a little bit.

31:46

So I found that on certain of our structured subtasks,

31:49

it actually exceeded, like I said, a compact frontier baseline.

31:57

So this is why routing policies have to be

31:59

re-evaluated constantly, too.

32:01

You have your baseline, you have your one that has

32:03

been winning, right?

32:04

They change these models all the time.

32:07

You have to constantly be updating and constantly be

32:10

re-evaluating what's happening.

32:11

The open models are getting better and better all the time.

32:14

And yeah, right now they're not at the point where we can replace

32:19

our main heavy reasoning workload.

32:23

But they're getting closer and ultimately it comes a

32:25

time when then you got to ask yourself whether you might want

32:28

to push it the rest of the way.

32:36

We also have this idea that, you know, when we're evaluating things,

32:39

right, if we want to get a little bit of a deeper understanding,

32:42

if maybe where one model performs better, if I want to get more

32:46

granular than just this task, this model, now I want to look at,

32:50

let's say that there are multiple models that are fairly equivalent.

32:53

When it comes to the reasoning task, are there some things that

32:56

one does better and one does worse?

32:57

How do I dive into that a little bit, right?

33:00

Can I find out, for example, that one model performs really

33:04

well on C, the other one performs really well on TypeScript?

33:09

One works really well on large diffs, one not so much.

33:14

Maybe on smaller, less complex PRs, I can use a slightly smaller model.

33:20

And it does just as good a job getting the answer quicker.

33:24

So there's a lot of variables that are in there.

33:25

So once you get that initial harness set up and you understand

33:29

where these things are doing well and they're not, you can start

33:32

to get a little bit more granular in the way that you approach

33:35

the problem, especially as you understand your own space better.

33:39

This being code review, us all being engineers, we understand

33:44

that space fairly well.

33:45

We've all lived it for a long time.

33:49

Some people know C very well and are going to be very harsh when

33:52

it comes to their feedback, and other people not so much, right?

33:57

So some of the models perform better in different

34:00

circumstances, language, framework, complexity, repository size.

34:06

And in some cases, the latency is the thing, right?

34:08

So if you're in an online setting where say you're in the CLI or a

34:11

VS Code extension, it could be that in that moment, you wanna sacrifice

34:17

a little bit of your quality for latency, but you should at least

34:21

make that as an informed choice.

34:23

You should know that you're doing that and why you're doing it.

34:31

So this is what I was sort of alluding to is at a certain

34:34

point you get the frontier models doing extremely well

34:38

at some tasks and you start to see the open models creeping up for

34:42

some of those heavier tasks even.

34:45

And as they get to a certain point where they're, they're starting to,

34:48

they starting to get good enough.

34:50

That you're starting to wonder, okay, maybe I can push it

34:53

the rest of the way.

34:54

Maybe I can leverage the data that I have and sort of the

34:57

domain expertise that I have and I can use some fine-tuning

35:01

or RL on that task, especially if you have access to the data,

35:07

to make that model into something that's going to work extremely

35:10

well for your particular use case.

35:13

In a lot of cases, a lot of verticals, this actually might

35:16

make sense right away.

35:17

In some cases, right, where you don't have the frontier

35:20

labs chasing down your vertical, it does make a lot of sense

35:24

for code generation.

35:25

Obviously, that becomes a little bit harder because they have

35:28

a lot of money being pushed into that particular space right now.

35:33

But code review is a little bit different.

35:35

It requires very, very heavy, multi-step reasoning.

35:39

And some of the tasks, even, I haven't seen models solve yet.

35:44

Some of the bugs that I myself have in what I call my hard set

35:48

have not been able to solve them.

35:50

And that's ones where you have to make multiple assumptions and

35:52

then reason through multiple steps and then you get to the conclusion.

35:55

Those ones are very difficult and ultimately, when the models

36:00

get there, I'll be very happy.

36:02

Those don't happen a lot in the real world, so it doesn't

36:04

necessarily impact my real-world evaluations of how I look

36:08

at things, but I'm always curious to see, as these models

36:11

get better and better, how good their multi-step, multi-hop

36:14

reasoning ability becomes.

36:16

So I have that mostly for my own understanding of the

36:19

raw potential of these models.

36:23

But if anything you can take away from this talk, it's

36:25

basically that you have to approach the idea of whatever

36:29

problem you're trying to solve

36:31

If you approach it from this fundamental principles of I'm going

36:34

to do it using empirical data, I'm going to approach it as an

36:37

engineering problem that I'm going to measure and be able to map that

36:42

measurement into things that my business and my users care about.

36:46

That shows me that when this offline benchmark goes well for me,

36:52

My users like my product better.

36:54

They use it more.

36:55

They convert better, or any other metric that I care about.

36:59

If you set that up, then that's when your ability to make that

37:03

rapid progress is gonna take off.

37:06

All right, thank you.

37:14

Thank you so much to David.

37:17

Another round of applause for David, let's go.

37:23

Okay, so we talked a little bit about, you know, maybe not needing

37:27

to specialize your model until your evals are set up, until you

37:31

know kind of exactly what you need it to be the best in the world at.

37:35

And of course, then that begs the question, what happens

37:38

once I do have those evals, right?

37:41

What happens when I am ready to make my model the best

37:45

in the world at whatever

38:09

Similar nebulous object, and that's because I think a lot

38:11

of us in the room know that's what it is right now, right?

38:15

Agents, like models are in agents, are in systems of models.

38:21

It is, you know, our application stack has

38:24

become pretty, pretty complex.

38:28

And the model architectures themselves have

38:30

become very complex.

38:31

One of the advantages that the closed models have versus

38:37

the open community, we don't share the same advantage, right, is that

38:42

we are cooking all the time when it comes to model architecture.

38:46

We are doing all kinds of very fun things that make our models faster.

38:51

and smarter, and work at longer contexts more efficiently,

38:54

and make them more able to handle, you know, out of distribution data.

38:58

We were doing all kinds of fun things at an architectural

39:01

level with our models that make them more and more difficult

39:06

to customize with a standardized customization stack, right?

39:11

And so what that means is that when I want to change my model up,

39:15

When I want to teach a new trick or make it the best in the world

39:18

at my use case, you know, no longer are the days where I can be very

39:23

confident that the tool is going to work out of the box day zero with

39:27

this brand new model because we're not all training the same model

39:31

Back in the days where we were all training different

39:34

versions of Llama, that was great.

39:36

But in the current age, that's not the case.

39:40

So this is just one of the complexities.

39:42

Another complexity, of course, is what David talked to us

39:45

about, which is evaluation.

39:48

It's very difficult to evaluate models.

39:51

It's really, really difficult to evaluate a system that

39:55

can do PhD-level, you know, I keep saying that phrase, but it's

40:00

a banger, PhD-level activities, but also, like, can't count, right?

40:05

It's very difficult to measure a system that is, in some

40:08

areas, otherworldly intelligent.

40:11

And then in others, you know, maybe not so much.

40:16

So even figuring out what to evaluate is difficult.

40:20

We've been trying to evaluate each other for a very long time and

40:24

we're still very bad at it, right?

40:26

Think of like interview processes.

40:29

Think about how do you evaluate someone's capabilities

40:31

when you meet them?

40:32

You know, we can't just probe them with a series of 150

40:36

math questions and say, oh, I bet you're a great guy, right?

40:40

The model's capabilities are growing to the point that

40:43

even just figuring out what to evaluate is becoming difficult.

40:48

With smaller models we have some advantages and with open

40:52

models especially smaller open models we have big advantages

40:55

which is that they have usually very limited narrow scope

41:00

and that makes them a dream to evaluate right if we need our

41:04

model to be excellent at calling a specific set of tools that

41:08

we know we need to be good at so that we can be good at our use case

41:12

It's not so bad, right? We can build an evaluation for that.

41:16

However, maybe not so much for everything else.

41:18

So the number of things you can do, of course, unit eval.

41:22

So this is like your classic exact match, right?

41:25

I need this JSON blob.

41:26

Did you give me this JSON blob?

41:29

LM is a judge.

41:30

We get a little bit more complicated.

41:31

We start to look at things maybe like sentiment.

41:33

We start to look at things like, you know, is this...

41:37

directionally correct, you know, is this verbose, is

41:40

this, you know, a pleasant answer, a human eval, obviously,

41:44

this is a extremely good way to evaluate anything, get people

41:49

to, to tell you if it's doing good or tell you if it's doing bad.

41:53

And then of course, production signals, nothing beats people

41:58

telling you that your thing is not good at the task to make you

42:01

believe that you need to improve your model at that task, right?

42:06

We have a number of different ways we can fine tune using

42:10

the term a little bit loosely here, but forgive me if you will.

42:14

The first one is something that David talked about as

42:17

well, which is we can wiggle the prompt around, right?

42:21

There's lots of ways to do this.

42:24

Manually, which is we literally sit at the keyboard and we

42:27

type out prompts and we see what does better.

42:29

And there are loads of ways to do this a little bit more smart,

42:32

a little bit more like engineers.

42:33

DSPy and other frameworks can help us actually select

42:37

the best prompts on sets of tasks.

42:39

However, a lot of the time, these things all roll back into

42:42

we need that good eval, which is why it's so important to have that.

42:47

Of course, next we have good ol' SFT, right?

42:50

We're still doing it.

42:51

2026, we're still doing SFT.

42:54

Maybe we're doing less of it now than we used to.

42:56

We still, maybe it's just part of like a warmup, right?

42:59

Or a cold start.

43:01

But the idea is that SFT in large part is still with us.

43:06

Show the model a lot of examples of what you want it to be good at.

43:09

And we're good to go.

43:10

One of the interesting tensions of the modern training stack,

43:14

right, is that we not only have the ability to do SFT, but

43:19

we also sometimes can get rid of the data shortage problem, right?

43:24

Which is the idea that with SFT, you kind of need a big

43:27

old pile of data.

43:28

We have now models that are so good at generating synthetic data

43:32

that we can kind of bridge the gap.

43:34

So SFT remains a strong contender, even in the kind of age of RL.

43:40

LoRa, classic, it's SFT but cheap.

43:43

There are tons of variants of LoRa.

43:46

There are tons of ways we can apply this technology and

43:48

there are a ton of benefits that we can leverage, especially through

43:51

the open model ecosystem when it comes to how we deploy the model.

43:55

Classic pattern still remains to train different LoRa heads

43:58

for different users or different functions and hot swap those

44:02

at inference time, right?

44:04

That's something that remains tried and true best practice even today.

44:10

And then finally,

44:12

You know, we are RL whatevers, right?

44:15

So RL is having a big moment right now, and for good reason,

44:20

especially now that these models have baked into them a

44:24

certain level of capability, right?

44:27

There's a virtuous cycle that I think we don't...

44:29

And then that is a virtuous cycle with how easily we can apply

44:33

techniques like RLHF, RLVR using things like GRPO or whatever other.

44:52

of the myriad, the zoo of RL techniques that we have today.

44:56

The big idea here, though, is that we have to know when

45:00

to choose which of these.

45:02

And the idea really bakes down a lot of the time to

45:06

external constraints.

45:07

So this is like, how much hardware do you have?

45:10

How much data do you have?

45:11

How much are you willing to spend on data or hardware, right?

45:15

If I have infinite piles of data and infinite GPUs, why not

45:20

just use SFT all the time, right?

45:21

But maybe you don't.

45:24

We certainly don't.

45:25

I certainly don't. There you go, okay.

45:28

When should we specialize?

45:30

When should we not? David touched on this a little bit,

45:32

but fine tuning is great when we're doing something like we have a

45:36

desired behavior or output, right?

45:39

This is a class, this hasn't changed since 2022.

45:43

If you need it to output a specific JSON blob, fine tune it.

45:47

It'll do it, there you go, good model.

45:50

Skip it when though you don't have a lot of samples.

45:52

We don't want to overfit on like our 16 samples and then

45:55

we can't generalize to anything.

45:58

Use fine tuning when we have, you know, many new words,

46:04

okay, many new words that we need to teach the model.

46:08

The model doesn't know all of the legalese, right, out

46:11

of the box most of the time.

46:13

It certainly doesn't know how your company uses words.

46:16

Something like fine tuning can help here.

46:19

You know, when the difficulty that you're having is that the model

46:23

is just not smart enough, SFT or fine-tuning generally might not be

46:28

the tool that you reach for, right?

46:30

We can, with a big effort, you can get the model to become smarter,

46:34

but if it's missing that kind of foundational level of knowledge,

46:38

you're going to be Sisyphus with this boulder, and it's

46:42

going to keep falling back on you.

46:44

uh you know oh yeah you can read the rest of the slide

46:46

you get the idea uh you know if you have many samples of high quality

46:50

data it's great uh if you haven't built any evals yet probably it's

46:54

not great uh the big one though is when you're doing the same

46:58

thing over and over and over again when you have the same task over

47:02

and over and over again when it's

47:05

Week 3 of your model doing something, and it's done the

47:07

same task 17 million times, this is a great time to fine-tune.

47:13

This is a great signal that your model is going to keep

47:15

doing this, and because it's going to keep doing this,

47:18

maybe you can do something like get a smaller model in there.

47:21

Maybe you could do something like fine-tune a faster model to sit

47:26

in that process in order to do just that one task over and over again.

47:32

And please don't fine-tune if you haven't tried changing the prompt.

47:36

If you've done no exploration with your prompt...

47:40

Yeah, I'm sure we all know, please don't fine-tune, right?

47:44

You're, especially with modern LLMs who are trained to be

47:49

great at instruction following, a lot of effort can be put,

47:54

or can, or sorry.

47:56

You can spend a lot of time changing the prompt and you're

47:59

going to get a lot more value than if you go straight to tuning.

48:03

To talk about some of these cases, you know, when and

48:08

why to fine tune a model, we're actually going to pass

48:10

it off to Shraddha Sridhar, who's going to talk to us about something

48:15

that she and her team have been cooking up here at Team Green.

48:29

Hi, everybody. My name is Shraddha Sridhar.

48:32

I'm the product lead for AI for Chip Design at NVIDIA.

48:36

Today, I'm going to talk to you about a fascinating domain.

48:40

You've heard about picking the right models and fine-tuning

48:47

them, or perhaps not, but then there is consideration of

48:53

when you apply a model to a domain.

48:55

I'm here to talk about what happens when you deploy models to a

48:59

complex domain, and chip design is the most extreme version of that.

49:05

Let's start by understanding why chip design is so hard

49:08

and how some of the learnings that we learned within chip design

49:12

will apply to any complex domain.

49:15

Our feedback cycles are long, 18 months.

49:20

You know, you ship a chip and then you find out whether the

49:23

decisions you've made were right.

49:25

There is no A-B testing, no rollback, only respins.

49:30

Speaking of which.

49:32

If you make a mistake and the bug makes its way to silicon, you

49:36

end up spending more than hundreds of millions of dollars, not

49:40

to mention the time wasted, right?

49:42

So given the high stakes, it's super important that

49:46

we prioritize correctness.

49:48

Our hardware engineers are highly trained to prioritize correctness.

49:53

They will not tolerate an answer that merely looks right.

49:57

They need to be able to verify because they cannot

50:00

afford to, right?

50:02

The chip design process itself is also pretty unique.

50:05

You have 10 stages or more going from deciding to build something

50:12

and writing a specification for it to taping it out, which is

50:16

sending the chip for fabrication.

50:18

And each of these stages are highly specialized.

50:22

You have a dedicated team of hardware engineers focusing

50:26

on that stage alone, and they have honed their intuition for that

50:31

stage over years, if not decades.

50:35

And that is the context in which we have increasing complexity

50:39

in architectures, so there is always parallel work in

50:45

each of these stages, creating a lot of coordination drag

50:48

among all of our engineers, and that's the setup in which we are

50:52

introducing probabilistic outcomes.

50:56

What all of that means is the goal is not to enable

51:00

our individual engineers go faster in their specific stages,

51:04

but rather enabling them to go across all the stages and

51:09

expedite chip design so that we can build chips faster with higher

51:14

quality and more generational gains going from one chip to another.

51:18

I'll talk about the lessons we learned as we work

51:22

towards this goal.

51:26

So the patterns I'm about to share are from production.

51:29

We have over 100 agents deployed in different engineering

51:33

workflows, and 10,000 engineers are using our agents.

51:38

They've made collectively a million queries, but the

51:41

path here was not a straight line.

51:44

We did learn, we did fail, we succeeded, and now we have

51:48

a clarity on the path forward.

51:52

So what did we try first?

51:55

In 2023, our team launched Chip NeMo, which are domain

52:00

adapted LLMs for chip design.

52:02

The research was strong, benchmarks were great, but then we learned

52:07

certain things in the production that we were not prepared

52:10

for and didn't anticipate through benchmarks.

52:14

The first was around data quality.

52:17

Now, a central ML team picked data for curation for training,

52:22

whereas domain intuition lived with the hardware engineers.

52:26

They alone were able to say what good looked like.

52:30

So that gap meant we couldn't provide quality responses.

52:35

And then we had the hardware engineers themselves who

52:37

prioritized correctness, so they were not ready for

52:40

black box answers yet.

52:42

They needed to verify the response.

52:44

And that's something that domain adaptive pre-training

52:48

is not best positioned to provide.

52:50

And also, the responses themselves were stale, because documents

52:54

and artifacts were developing faster than we could train models.

52:58

So the model was always playing a catch-up.

53:01

And all of this meant we didn't see adoption in production.

53:05

So that's the lesson we want to share with you all.

53:08

Think hard before you try pre-training, let alone

53:12

post-training and fine-tuning that David and Chris talked about.

53:16

Pre-training is another level.

53:18

You want to be super intentional when you attempt it at all.

53:22

So given lack of adoption, we pivoted and to a much

53:27

simpler solution.

53:30

So we said, we'll start with RAG.

53:34

Teams curate documents that apply to themselves because

53:37

they know what good looks like.

53:39

We put control in the hands of hardware engineering teams

53:42

and got them to not only curate documents, but also update them.

53:47

And they were responsible for collecting benchmarks,

53:50

running evaluation, so sort of controlling their own destiny.

53:54

We also ensured that responses are traceable to the documents

54:00

where they originated from.

54:03

So now we fixed the problem of traceability and visibility,

54:06

which meant engineers could trust the responses.

54:10

And that was the key to driving adoption.

54:13

Now, responses were also always fresh because it was a config

54:18

change at max to update documents.

54:20

There was also continuous indexing in the background,

54:23

so all of this meant the responses were fresh, they were accurate,

54:26

they were believable.

54:28

And the result was adoption.

54:32

Collectively, our engineers made more than a million queries,

54:35

and we also learned that they were not just answering questions

54:40

using these chatbots, they were also starting to build agents.

54:44

They were deploying these chatbots as APIs within their

54:47

hardware engineering workflows, and they were building agents.

54:50

So that's when we knew that we are at the next inflection

54:54

point, which is agents.

54:58

So now that we knew that agents were getting built, the next

55:01

step was to understand whether it mattered, right?

55:05

Are we really accelerating workflows or are we

55:08

just building agents?

55:10

So what we needed was an impact quantification framework.

55:14

We looked at all the options out there in terms of measuring ROI

55:19

and we found all of them lacking.

55:21

So what we did was we went back to the drawing board

55:24

and thought about what kind of foundation do we want to establish

55:28

to be able to measure impact.

55:31

We combined org design, product metrics, and infra hooks together

55:37

into one foundation.

55:38

We started by interviewing different teams and understood

55:43

what their processes were, where they were applying AI, and what

55:46

kind of value they were unlocking.

55:49

Once we had clarity on that, what we next did was we codified

55:53

it into our system.

55:55

So our developers could go to the system with a few lines of config,

56:00

talk about the value that the agent is unlocking for invocation,

56:06

and everything else was automated.

56:08

From there, the receipts came in.

56:10

And here is the interesting part.

56:12

We were surprised about one pattern, and another pattern

56:16

just confirmed our assumptions.

56:18

The first one was Root Cause Analysis.

56:21

This is a step that's performed across all aspects of chip design.

56:26

And it is the step that takes the most amount of time.

56:30

Like, for example, root cause analysis is a very common

56:34

task within design verification, and that takes about more

56:37

than 50% of the capacity of Chip Design development.

56:41

So it was an important step, and agents were getting built

56:45

to automate that step.

56:47

Now, the engineer or the developer reported that every time the

56:52

agent was invoked, it was saving them two hours.

56:56

So obviously this is what I was expecting and it was

56:59

heading in the right direction in terms of productivity gains.

57:03

The second pattern though was really interesting and

57:06

something that surprised me.

57:08

This was not at all about time savings.

57:10

Here, the developer said, from a power optimization agent, they

57:16

were able to unlock Power Insights three full months ahead of tape

57:22

out, which meant any optimization insights that they unlocked, they

57:26

were able to action on it within the same generation of the chip.

57:32

Before AI, these insights would arrive too close to tape out, and

57:37

they arrived in spreadsheets and PowerPoint presentations, and they

57:43

were hard for engineers to trust.

57:45

So they mostly did the function of sign-offs, but as now, with all

57:50

these additional insights unlocked through AI, they were able to

57:54

take advantage of these additional optimizations, which meant

57:56

we were pushing the boundaries of performance within the same chip.

58:00

That's the step in the right direction.

58:02

So we want to understand if this was a one-off situation or

58:06

if there was a pattern behind it.

58:10

So what we looked at was among all the 100 agents that were getting

58:14

built, what kind of patterns we saw in terms of adoption.

58:17

So the first one is where everything, most agents fell,

58:21

which was a great curve.

58:23

And these are individual productivity agents.

58:25

You adopt them. It's very easy to adopt.

58:28

You quickly see the value and you unlock

58:31

productivity gains out of it.

58:33

The next one was the team level aspect of productivity as well.

58:39

These are the same individual productivity agents deployed to the

58:42

whole team that's similar to the root cause analysis that we saw.

58:47

So in both of these cases, you're going to save the same

58:50

amount of time every time you invoke the agent.

58:53

So there is some value, but eventually it plateaus.

58:56

And that's what you would see in terms of the usage of these agents.

59:01

One of the interesting curves is the purple one, where adoption

59:05

is really slow, right?

59:07

Because the engineer takes time to trust the capability of the

59:12

agent, and even to understand how to make use of it in novel ways.

59:16

And this is about capability expansion.

59:19

This is way beyond productivity.

59:21

And this is the same pattern as the power optimization

59:24

agent that we saw.

59:25

But this, the purple curve, is at the individual engineer level.

59:30

When you translate that to the team level, then

59:33

you unlock new capabilities across the entire team.

59:37

So an agent that is used by one team and they end up using

59:42

it to increase capabilities actually compounds the value for

59:47

the entire Chip Design lifecycle.

59:49

So that's the pattern that we wanted to double down on.

59:53

Now that we knew what capability expansion looks like, the next step

59:57

was for us to understand how to build intentionally towards this.

01:00:02

So we went back to a framework.

01:00:05

So we looked at what agents were picking up with their

01:00:09

teams, and we looked at the

01:00:13

production data and user insights.

01:00:16

From there, we realized that we really do need the task

01:00:18

tools to quickly build trust and see time to value.

01:00:23

From there, it gave us valuable insights on how to build team-wide

01:00:27

context and extend the individual productivity to the team level.

01:00:33

And that's where you hit the connective tissue sort of agents.

01:00:36

So these are all agents that save time for the entire team.

01:00:41

And then we also have power tools, which is where you're expanding

01:00:45

capability at the individual level.

01:00:49

So all of this data, when you launch in production, you get a

01:00:52

sense for where the team is headed.

01:00:55

So you're actually moving along with the team and

01:00:57

figuring out how to...

01:01:00

Set up this virtual cycle by using shared context and

01:01:04

also expanded capabilities.

01:01:06

And now combine them and you end up in the flywheel zone.

01:01:10

And this is the zone you want to be in as you improve productivity

01:01:15

and as you also use AI to unlock new capabilities.

01:01:19

The extreme version of it is what I want to show you in a vision demo.

01:01:23

But before that, we'll also talk about some of the optimization

01:01:30

opportunities that we identified.

01:01:34

So far, we've seen that we started simple with RAG and

01:01:38

then worked our way towards capability expansion engines.

01:01:41

And the key learning for us over there has been

01:01:44

Start simple because you do need to build trust and then eventually

01:01:48

earn your way to optimization.

01:01:50

But once you do have real patterns that stick for optimization

01:01:55

and you're confident that you have production data,

01:01:58

That is proprietary.

01:01:59

You are now ready to optimize.

01:02:01

Like in our case, we picked a model that's open source model

01:02:06

and fine-tuned on proprietary data.

01:02:09

And you can see that the output on CVDP, which is

01:02:14

a hardware-specific benchmark, is comparable to SODA models.

01:02:19

Then if you combine that with cost and latency benefits,

01:02:23

it becomes a compelling alternative for those hardware-specific tasks.

01:02:28

But then, that is only part of the story.

01:02:31

When you put that model in a loop, agentic loop, and combine

01:02:36

frontier model and fine-tune model, that's when the magic happens.

01:02:41

Previously, you saw that

01:02:43

The capability for code modification was at 67%,

01:02:47

but combining frontier models with fine-tuned models, now we were able

01:02:51

to achieve more than 90% across most of the code generation tasks.

01:02:56

That's what you can achieve when you close the loop and

01:02:59

you see compounding.

01:03:02

So that is the entire arc.

01:03:03

We started with simple RAC scenarios, then we looked

01:03:07

at agents that were getting built, and we graduated towards

01:03:10

cross-team agents.

01:03:11

And from there, we find optimization opportunities

01:03:15

to further compound benefits.

01:03:18

So now, let's switch gears to talk about what the vision

01:03:21

looks like for this area.

01:03:26

So this is the vision.

01:03:27

We talked a little bit about how we can improve capability.

01:03:34

of an agent by identifying production uses.

01:03:39

From there, as these agents are getting used and engineers

01:03:42

find novel ways to apply them, their agency improves and they are

01:03:47

more incentivized to use the model.

01:03:49

So that's what we mean by the frontier frywheel.

01:03:52

So this is the vision.

01:03:54

The engineer adopts the AI system and uses it.

01:03:58

So that gives us implicit feedback on what areas to improve for.

01:04:04

Then, the engineer, being the domain expert, also corrects the

01:04:09

system, not only when it is wrong, but also when the assumptions

01:04:15

are subtle and are violated.

01:04:19

Between the corrections and the corrected assumptions,

01:04:23

we then capture all of that data programmatically and

01:04:27

use that for training and also for improving the system.

01:04:31

So as this process compounds, you slowly expand the capability

01:04:37

of the system entirely, and eventually you work towards

01:04:40

the center of the flywheel, which is domain generalists.

01:04:45

Today, we do not have hardware engineers that go across the

01:04:48

entire lifecycle, but this is a mechanism that could

01:04:52

potentially take us there.

01:04:54

And then when you do that, you're tackling the bottlenecks

01:04:58

that currently exist in the current assembly line sort

01:05:01

of process of chip design.

01:05:04

So that's the goal.

01:05:05

We start small, we have our engineers adopt the system,

01:05:10

we learn from them, and then we codify that into the system.

01:05:15

I have a short demo to show you how one rotation or revolution

01:05:20

of this flywheel looks like.

01:05:33

What are the spots?

01:05:52

So Sharada, I think they're playing it from back there,

01:05:54

so can we switch to the on-camera device, on the podium device?

01:06:21

And if not, we might have to just show them later.

01:06:34

What do you say?

01:06:36

Come on.

01:06:39

Big round of applause.

01:06:44

Awesome. This will be worth it, I'm telling you.

01:06:54

We like to keep people on their toes.

01:06:58

Thank you all for your patience.

01:06:59

So here is the system.

01:07:00

This is a vision prototype, and this is just to elucidate

01:07:04

how the flywheel will rotate and get to one revolution of it.

01:07:11

So this engineer, who is an expert now across the entire Chip

01:07:18

Design lifecycle, uses a system.

01:07:21

And what they do is they write a design spec for a new feature

01:07:26

and run the full pipeline.

01:07:37

Okay, come on.

01:07:46

Try again.

01:07:50

There you go. So it's running through architectural

01:07:53

specification, article generation, test bench, formal

01:07:57

verification, and PPA analysis.

01:07:59

Eventually, it's generating a DB report.

01:08:04

And what happens is the model violates some subtle

01:08:07

assumptions, and it fails.

01:08:12

So there's a bug detected now.

01:08:14

Formal verification stage flagged a bug.

01:08:17

So here is where the export comes into the loop.

01:08:23

They look at the bug,

01:08:34

They understand the model's reasoning and they see where

01:08:38

the model went wrong and then they enter the corrected code.

01:08:44

Not just that, they talk about why the model got that bug wrong,

01:08:50

which means the next time the model has to make any sort of decision

01:08:54

that resembles this feature, it's going to get it right.

01:09:01

The engineer also provides the reason for the fix.

01:09:04

So all of this becomes training data for the

01:09:06

model to continuously learn.

01:09:09

So now we save and learn.

01:09:13

And note what happens to the CVDP score here.

01:09:15

So I'm gonna run the full pipeline.

01:09:18

So this is again, the system that is autonomously running

01:09:21

through all the stages.

01:09:23

And we note the CVDP score.

01:09:28

So in this case, the model improved because it got valuable feedback

01:09:33

from that domain expert that no model simulation alone can provide.

01:09:39

And here is the process.

01:09:41

It's eventually, you want to reach a certain level to

01:09:45

be tape-out ready.

01:09:46

And simple RL alone, from a model standpoint, would

01:09:51

not take you to the extent that a real-life expert's feedback would.

01:09:57

And that's what you see over here.

01:09:59

The expert corrects and keeps pushing the boundary of what

01:10:04

the model can perform over time.

01:10:08

So that's the concept.

01:10:09

We have all of this training signal that we are capturing

01:10:12

from real domain experts who've honed their intuition over years.

01:10:19

So that is the vision.

01:10:20

This is domain experts steering AI to build chips based on

01:10:29

real-life AI workloads.

01:10:31

So essentially, this is AI accelerating chip design to

01:10:35

build chips that accelerate AI.

01:10:38

And that is the vision.

01:10:39

At NVIDIA, we are lucky because we have access to the best

01:10:43

hardware engineers on Earth.

01:10:45

So we want to build our models and systems

01:10:48

So that we take advantage of that expertise and we further solidify

01:10:52

it with the expert in the loop.

01:10:55

So this is not about replacing humans, it is about augmenting

01:10:58

them and ensuring that we are taking everybody along.

01:11:03

That's the vision.

01:11:04

So as I help you, thank you.

01:11:09

Thank you, everybody.

01:11:11

So as I look at how you're gonna be building AI for your

01:11:15

domain, especially if it is complex, I wanna leave you

01:11:18

with three questions.

01:11:22

And now can we swap back to the slides?

01:11:31

It's OK, I can talk through it.

01:11:34

Oh, amazing, thank you.

01:11:37

So how are you building trust with your users?

01:11:40

Can your users directly interrogate AI?

01:11:43

That is step number one, because trust comes before adoption.

01:11:47

And then, how are you capturing expert feedback?

01:11:51

And remember, we are talking about a complex domain here, so feedback

01:11:55

is not otherwise available.

01:11:56

It's all in the heads of your experts.

01:11:58

Everything is undocumented.

01:11:59

So how are you capturing them, and how are you making use of that

01:12:02

feedback to improve your system?

01:12:03

So that's question number two.

01:12:06

And the third question is, what do you uniquely have

01:12:10

from months of production that nobody else can buy?

01:12:14

These questions took us months to figure out and

01:12:18

come up with an answer.

01:12:20

I hope they save you some time.

01:12:22

Thank you all.

01:12:34

All right, and you got me to close this out.

01:12:36

So, excellent stuff.

01:12:38

Thank you so much, Shraddha, for that.

01:12:41

Very brave to do a live demo, guys.

01:12:43

You all know that. So another round of applause for Shraddha

01:12:45

for the live demo.

01:12:46

Let's go.

01:12:49

I'm going to talk very quickly about deployment, which is

01:12:52

something we want to make our models also produce tokens quickly.

01:12:58

Deployment, just like customization, is something

01:13:01

that, again, the Frontier model shops have a very

01:13:06

specific advantage for, which is that they are building their model.

01:13:09

That's all they do. They build one model, and they build

01:13:12

it good and fast.

01:13:14

Us in the community using open models have to contend with

01:13:18

architectures that are designed to make our models go faster,

01:13:23

but are also non-standard, right?

01:13:25

So we have models that can just chew through tokens.

01:13:30

But the decisions that, say, NVIDIA made for Nemotron 3 Nano and

01:13:34

Super and, you know, Ultra coming up, is not the same decisions

01:13:38

that someone like Kimi makes.

01:13:40

The idea is that we have this kind of, you know, large gap

01:13:50

in architectures, and we need to have very, very well-tuned

01:13:55

inference backends to deploy these things well, and that

01:13:59

means that it's a lot of effort to make them work across the board.

01:14:03

We also have this real

01:14:05

This real thing that's happening, which is that we are moving

01:14:09

from this world of dense models to this world of MoEs.

01:14:14

MoEs have not much more burden for deployment, but definitely

01:14:22

deploying an MoE extraordinarily effectively is more difficult

01:14:25

than the same for a dense model, for a number of reasons.

01:14:30

Dense models are simple.

01:14:31

We can just stand them up, no probs, with default

01:14:34

configs from a lot of the inference backends, right?

01:14:39

We get very predictable latency.

01:14:41

We get very predictable fine-tuning like we talked about earlier.

01:14:44

We can even have one, you know, a model running a 100

01:14:48

billion parameter model, or a GPU running a 100 billion

01:14:52

parameter model, one GPU, right?

01:14:54

That's crazy.

01:14:55

Every parameter is active for every forward pass.

01:14:58

We get less flop efficiency, but that's okay.

01:15:02

And then memory-bound at large batch sizes.

01:15:06

You guys know some of this stuff.

01:15:07

The idea is that dense models have a lot of really

01:15:11

interesting advantages.

01:15:13

Bars MOE models, on the other hand, have a few.

01:15:16

We get only active parameters are computed per token.

01:15:20

So this is great, right?

01:15:21

So we see that nomenclature of like 120B, and then 12AB, right?

01:15:28

Which means that even though our model's taken up 120 billion

01:15:32

parameters worth of space, we only actually care about 12 when

01:15:35

we're doing our forward passes.

01:15:37

Then we have, of course, higher capability at low of,

01:15:41

Lower active flop cost.

01:15:44

This is huge, right?

01:15:45

This means that we can borrow intelligence from

01:15:49

the model in a way.

01:15:51

The reason that's important is because it helps us to

01:15:54

make sure that we're not, you know, we're leveraging

01:15:57

only the part of the model we need, but during training, when we're

01:16:00

making these things, they're...

01:16:03

They're being trained across a broad variety of topics.

01:16:06

And so we can borrow from experts, right?

01:16:09

If we have our task that our model has been trained at, has been

01:16:12

trained very well at, well, when we do some task that's slightly out of

01:16:15

distribution, we can borrow from an expert that might've been trained

01:16:20

closer to that distribution.

01:16:22

However, it's tough to pack the model into the GPUs.

01:16:28

Of course, we have a lot of different ways we can serve.

01:16:31

I've put the four, you know,

01:16:35

biggest, most popular, I don't know how you define it, but

01:16:38

the idea is like, these are the ones you're gonna see first

01:16:41

if you're trying to deploy models.

01:16:45

A little bit biased, but NVIDIA NIM obviously is a good solution for

01:16:49

picking the best of the next three for your use case based on the

01:16:53

incredible work of our engineers.

01:16:55

But then we have vLLM, open source standard.

01:16:58

It is the most mature ecosystem by far today.

01:17:02

It is typically, you can expect day zero support

01:17:06

in some fashion, right, for a lot of models that drop.

01:17:10

SGLang, of course, up and comer.

01:17:13

Great to use, great DevEx,

01:17:17

ecosystem is still being refined, and TensorRT-LLM, obviously,

01:17:24

you know, hopefully designed to be the best for our models,

01:17:27

something like the Nemotron 3 super launch, TensorRT-LLM is

01:17:32

going to get us the most number of tokens per user, per second, right?

01:17:37

This is the idea.

01:17:39

When we're choosing between these backends.

01:17:42

But the big idea here is that we are spoiled for choice.

01:17:45

And so when we choose a model to power our agents using open

01:17:50

weights models, we are also having to choose an inference backend.

01:17:54

and we haven't talked about it on a slide but we're also

01:17:56

choosing a customization framework and we're also choosing you know

01:18:00

what hardware we want to deploy it on we're also choosing we have

01:18:03

so many choices to make in 2026 and so uh we've got some key takeaways

01:18:08

here to uh to to kind of wrap us up and then we're gonna i think

01:18:12

we have some time for questions uh so we're gonna start with

01:18:16

evals before you start with models

01:18:18

Without domain evals, you're optimizing blind, classic,

01:18:22

slide, you know, there you go, David said it really well,

01:18:25

the idea is if you cannot test it, then you cannot measure

01:18:28

it, and if you can't measure it, what's the point, the hard decision

01:18:31

isn't just open versus frontier, it's knowing when you're use cache.

01:18:36

UseCase has matured to the point where you can make that choice.

01:18:40

I would not say, if you are just getting started with

01:18:43

a company, you're just starting your company today, please

01:18:47

don't spend two months waffling over, should I use a frontier

01:18:52

model here and an open model there?

01:18:56

Get to a point where you're able to make the decision

01:18:59

before you start making it.

01:19:01

If you're trying to fold in something like Nemotron 3

01:19:05

Nano or one of the smaller Quens, this is a much easier decision for

01:19:09

you to make than if you're trying to decide if you want to go with

01:19:12

something like Kimi K2.5, right?

01:19:15

They're entirely different engineering challenges.

01:19:19

Make sure that when you are choosing open models, you

01:19:22

are aware of those challenges.

01:19:24

You should just start with SFT or LoRa if you need to do some

01:19:29

light customization to your model.

01:19:31

You do not have to go deep into asynchronous rollouts

01:19:34

and RL to make your model a little bit better at the task.

01:19:38

Start with a hammer before you use a scalpel.

01:19:42

And please, please don't pre-train a model unless you really need to.

01:19:48

A lot of people.

01:19:49

Especially at a conference like GTC, we'll ask, you know,

01:19:52

should I pre-train my own model?

01:19:54

And the answer is, if you've asked me, should

01:19:55

I pre-train my own model?

01:19:56

You definitely shouldn't. Just pull a model off the shelf.

01:20:00

Companies like NVIDIA, by the way, release base model

01:20:03

checkpoints that you can start from if you want to start from

01:20:05

a place and customize from there.

01:20:08

And then, of course, your deployment stack matters,

01:20:10

and it is yet another choice.

01:20:13

Thank you so much for everybody's time, and we're going to do

01:20:15

some questions now.

01:20:17

all right why don't we bring the other presenters back

01:20:19

up on stage who's got a question for anybody who talked they

01:20:24

did they give you all the answers there are no questions oh way back

01:20:28

there okay somebody's gonna make me run oh solid job Thanks, Saunders.

01:20:38

Hi, a little bit hard to remember with all that content but

01:20:43

I have one question you said that a pivotal moment was

01:20:47

when you combine the frontier models and the local models and how

01:20:54

do you do that in your application?

01:20:57

Do you use that LLM router or what are you using that?

01:21:01

Is that for me?

01:21:02

Oh, yes, for you.

01:21:03

It's, I think, for the lady, I forgot the name, but you

01:21:07

said that, I think.

01:21:08

Yeah, absolutely.

01:21:09

So, we do have an example where we combined Frontier

01:21:12

Model, Cloud Sonnet 4, and

01:21:16

A fine-tuned model.

01:21:17

So what we did in here was we closed the agentic loop and

01:21:20

optimized around the capabilities of individual models.

01:21:24

So essentially, the fine-tuned model would generate code

01:21:28

that is hardware-specific, and then we had Cloud reflect on the...

01:21:35

A code that was generated and then it would coordinate

01:21:37

across all the outputs.

01:21:38

So essentially it was an agentic loop where the whole loop was

01:21:42

learning from itself and the model was also improving over time based

01:21:45

on the feedback it got from cloud.

01:21:48

And then on top of that, I know that we have talked about

01:21:51

our LLM router, which we have, which is meant to do some

01:21:56

of this work for you.

01:21:57

So if you're wondering how to stitch everything together,

01:22:01

we have an LLM router that you can check out on GitHub.

01:22:05

All right, and I think one more question.

01:22:07

I saw you raise your hand.

01:22:10

Hello, my name is Jonathan, and I'm a Gen.AI team lead.

01:22:14

And my team has always had this question.

01:22:16

How exactly do we even start designing a complex agentic

01:22:20

workflow with open source models?

01:22:23

Like, we've always had the problems with doing models, trying

01:22:26

to delegate to other models, and they just fail miserably at that.

01:22:30

Betrayed context.

01:22:32

We've tried so many different ways.

01:22:35

How can we even get started doing it the right way?

01:22:38

I think David's going to have a great first answer for this.

01:22:44

I think when you get started, you start with the smallest

01:22:46

part of the problem, right?

01:22:47

So what am I trying to accomplish?

01:22:49

If I'm building a large workflow, in my case, like a code review,

01:22:52

I think the first thing I would try is just getting

01:22:55

straight from the diff into the actual review product.

01:22:58

And if I'm thinking about the agentic loops for me, it's like

01:23:01

There's a couple different versions of it.

01:23:03

I'm not going to put the whole thing into

01:23:05

one gigantic agentic loop.

01:23:06

I'm going to see whether I can break it into pieces.

01:23:09

Which pieces can I break it into that then I can validate

01:23:11

that on a smaller scale before I then maybe add more tools

01:23:15

or add more capabilities.

01:23:16

So try and shrink your world a little bit first

01:23:19

and build it from there.

01:23:21

Even that portion of it, figuring out the prompt, bringing out

01:23:24

the strategic part of it, and getting that system to

01:23:26

work well, does take effort.

01:23:28

But once you've got the foundation and you've learned some of

01:23:30

the lessons along the way there, I think it becomes a little bit

01:23:33

easier to sort of tack things on.

01:23:36

And it kind of like, you know, say it the same way that we

01:23:39

explain to people who are like, how do I approach a

01:23:42

complex problem with an LLM, right?

01:23:44

It's all about task decomposition.

01:23:46

You want to decompose whatever your pipeline is into its atomic parts.

01:23:51

And if you understand each of those atomic parts, like

01:23:54

David just said, you're going to be able to understand, you know, which

01:23:57

is the one that I can pass off.

01:23:59

This atomic part is very important, but it doesn't need to think

01:24:03

a lot, or we would benefit from it being much faster.

01:24:07

And I think that's the idea.

01:24:09

The better you can understand exactly how your agent decomposes

01:24:13

and what are the atoms of work it's going to do, the

01:24:16

easier it's going to be to say, we've noticed that this atom of

01:24:20

work needs to be done 10,000 times per run, and we need it to be fast.

01:24:25

And that's where you can say, maybe we should try plugging

01:24:27

something open in.

# Agentic AI Using Open Source Models: Finetuning, Deployment and Chip Design Case Studies

Chris Alexiuk,Sr. Product Research Engineer,NVIDIA

David Loker,VP of AI,CodeRabbit

Shraddha Sridhar,Product Lead, AI for Chip Design,NVIDIA

AI-Generated Summary of this Video

**Highly Rated**

Rate Now

Open models have evolved from experimental alternatives to production-grade building blocks for agentic AI and enterprise applications. But deciding when to use open source models, and choosing and customizing them requires a clear framework. 

In this session, we'll walk through a practical decision framework for deploying open models at scale from picking the right model size for your agentic tasks to the software and practices for customizing, evaluating, and deploying them. We'll bring it to life with two case studies: how CodeRabbit runs open-source models in production, and how NVIDIA’s chip design team turned early failures into a repeatable flywheel: using a decision framework to pick use cases, then fine-tuning open-source models for the patterns that stuck. Whether you're building agentic AI, internal copilots, or cost-sensitive services, this talk will give you the framework and tools to ship with confidence.

### Learn More About This Topic

Share

Favorite

Add to list

PDF 

Events & Trainings:GTC San Jose

Date:March 2026

Topic:Agentic AI / Generative AI - Code / Software Generation

Industry:All Industries

Level:Technical – Advanced

Language:English

NVIDIA technology:NeMo, NVIDIA NIM

Region:

Company Information

*   [About Us](https://www.nvidia.com/en-us/about-nvidia/)
*   [Investors](https://investor.nvidia.com/home/default.aspx)
*   [Venture Capital (NVentures)](https://www.nvidia.com/en-us/startups/nventures/)
*   [NVIDIA Foundation](https://www.nvidia.com/en-us/foundation/)
*   [Research](https://www.nvidia.com/en-us/research/)
*   [Corporate Sustainability](https://www.nvidia.com/en-us/sustainability/)
*   [Technologies](https://www.nvidia.com/en-us/technologies/)
*   [Careers](https://www.nvidia.com/en-us/about-nvidia/careers/)

News and Events

*   [Newsroom](https://nvidianews.nvidia.com/)
*   [Company Blog](https://blogs.nvidia.com/)
*   [Technical Blog](https://developer.nvidia.com/blog/)
*   [Webinars](https://www.nvidia.com/en-us/about-nvidia/webinar-portal/)
*   [Stay Informed](https://www.nvidia.com/en-us/preferences/email-signup/)
*   [Events Calendar](https://www.nvidia.com/en-us/events/)
*   [GTC AI Conference](https://www.nvidia.com/gtc/events/)
*   [NVIDIA On-Demand](https://www.nvidia.com/en-us/on-demand/)

Popular Links

*   [Developers](https://developer.nvidia.com/)
*   [Partners](https://www.nvidia.com/en-us/about-nvidia/partners/)
*   [Executive Insights](https://www.nvidia.com/en-us/executive-insights/)
*   [Startups and VCs](https://www.nvidia.com/en-us/startups/)
*   [Documentation](https://docs.nvidia.com/)
*   [Technical Training](https://www.nvidia.com/en-us/learn/organizations/)
*   [Professional Services for Data Science](https://www.nvidia.com/en-us/support/enterprise/advisory-services/)

Follow NVIDIA 

[](https://www.facebook.com/NVIDIA "<util:I18n key=\"Follow GeForce on Facebook\" />")[](https://www.instagram.com/nvidia/?hl=en)[](https://www.linkedin.com/company/nvidia/)[](https://twitter.com/nvidia "<util:I18n key=\"Follow GeForce on Twitter\" />")[](https://www.youtube.com/user/nvidia)

[United States](https://www.nvidia.com/en-us/location-selector/)

*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [Your Privacy Choices](https://www.nvidia.com/en-us/about-nvidia/privacy-center/)
*   [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/)
*   [Accessibility](https://www.nvidia.com/en-us/about-nvidia/accessibility/)
*   [Corporate Policies](https://www.nvidia.com/en-us/about-nvidia/company-policies/)
*   [Product Security](https://www.nvidia.com/en-us/product-security/)
*   [Contact](https://www.nvidia.com/en-us/contact/)

Copyright © 2026 NVIDIA Corporation

Select Location

The Americas

*   [Argentina](https://www.nvidia.com/es-la/ "Argentina")
*   [Brasil (Brazil)](https://www.nvidia.com/pt-br/ "Brasil (Brazil)")
*   [Canada](https://www.nvidia.com/en-us/ "Canada")
*   [Chile](https://www.nvidia.com/es-la/ "Chile")
*   [Colombia](https://www.nvidia.com/es-la/ "Colombia")
*   [México (Mexico)](https://www.nvidia.com/es-la/ "México (Mexico)")
*   [Peru](https://www.nvidia.com/es-la/ "Peru")
*   [United States](https://www.nvidia.com/en-us/ "United States")

Europe

*   [België (Belgium)](https://www.nvidia.com/nl-nl/ "België (Belgium)")
*   [Belgique (Belgium)](https://www.nvidia.com/fr-be/ "Belgique (Belgium)")
*   [Česká Republika (Czech Republic)](https://www.nvidia.com/cs-cz/ "Česká Republika (Czech Republic)")
*   [Danmark (Denmark)](https://www.nvidia.com/da-dk/ "Danmark (Denmark)")
*   [Deutschland (Germany)](https://www.nvidia.com/de-de/ "Deutschland (Germany)")
*   [España (Spain)](https://www.nvidia.com/es-es/ "España (Spain)")
*   [France](https://www.nvidia.com/fr-fr/ "France")
*   [Italia (Italy)](https://www.nvidia.com/it-it/ "Italia (Italy)")
*   [Nederland (Netherlands)](https://www.nvidia.com/nl-nl/ "Nederland (Netherlands)")
*   [Norge (Norway)](https://www.nvidia.com/nb-no/ "Norge (Norway)")
*   [Österreich (Austria)](https://www.nvidia.com/de-at/ "Österreich (Austria)")
*   [Polska (Poland)](https://www.nvidia.com/pl-pl/ "Polska (Poland)")
*   [România (Romania)](https://www.nvidia.com/ro-ro/ "România (Romania)")
*   [Suomi (Finland)](https://www.nvidia.com/fi-fi/ "Suomi (Finland)")
*   [Sverige (Sweden)](https://www.nvidia.com/sv-se/ "Sverige (Sweden)")
*   [Türkiye (Turkey)](https://www.nvidia.com/tr-tr/ "Türkiye (Turkey)")
*   [United Kingdom](https://www.nvidia.com/en-gb/ "United Kingdom")
*   [Rest of Europe](https://www.nvidia.com/en-eu/ "Rest of Europe")

Asia

*   [Australia](https://www.nvidia.com/en-au/ "Australia")
*   [中国大陆 (Mainland China)](https://www.nvidia.com/zh-cn/ "中国大陆 (Mainland China)")
*   [India](https://www.nvidia.com/en-in/ "India")
*   [日本 (Japan)](https://www.nvidia.com/ja-jp/ "日本 (Japan)")
*   [대한민국 (South Korea)](https://www.nvidia.com/ko-kr/ "대한민국 (South Korea)")
*   [Singapore](https://www.nvidia.com/en-sg/ "Singapore")
*   [台灣 (Taiwan)](https://www.nvidia.com/zh-tw/ "台灣 (Taiwan)")

Middle East

*   [Middle East](https://www.nvidia.com/en-me/ "Middle East")

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/). You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control (GPC) signal and recorded your rejection of all optional cookies on this site for this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. To opt out of non-cookie personal information "sales" / "sharing" for targeted advertising purposes, please visit the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies. You can manage these settings in the NVIDIA [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies which overrides at least one of your previous settings. You can manage them in the [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

Manage Settings

Reject Optional Accept All

![Image 2: Company Logo](https://cdn.cookielaw.org/logos/10ddf4ca-c072-45d0-b3ac-eead0ed93db0/6e17f6e4-c77b-4a11-9f34-c107c42e4bfc/7981.png)

Cookie Settings

We and our third-party partners (including social media, advertising, and analytics partners) use cookies and other tracking technologies to collect, store, monitor, and process certain information about you when you visit our website. The information collected might relate to you, your preferences, or your device. We use that information to make the site work, analyze performance and traffic on our website, provide a more personalized web experience, and assist in our marketing efforts.

Under certain privacy laws, you have the right to direct us not to "sell" or "share" your personal information for targeted advertising. To opt-out of the "sale" and "sharing" of personal information through cookies, you must opt-out of optional cookies using the toggles below. To opt out of the "sale" and "sharing" of data collected by other means (e.g., online forms) you must also update your data sharing preferences through the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/).

Click on the different category headings below to find out more and change the settings according to your preference. You cannot opt out of Required Cookies as they are deployed to ensure the proper functioning of our website (such as prompting the cookie banner and remembering your settings, etc.). By clicking "Save and Accept" or "Decline All" at the bottom, you consent to the use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) in accordance with your settings and accept our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). For more information about our privacy practices, please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/).

Required Cookies

Always Active

These cookies enable core functionality such as security, network management, and accessibility. These cookies are required for the site to function and cannot be turned off.

Cookies Details

Performance Cookies

- [x] Performance Cookies 

These cookies are used to provide quantitative measures of our website visitors, such as the number of times you visit, time on page, your mouse movements, scrolling, clicks and keystroke activity on the websites; other browsing, search, or product research behavior; and what brought you to our site. These cookies may store a unique ID so that our system will remember you when you return. Information collected with these cookies is used to measure and find ways to improve website performance.

Cookies Details

Personalization Cookies

- [x] Personalization Cookies 

These cookies collect data about how you have interacted with our website to help us improve your web experience, such as which pages you have visited. These cookies may store a unique ID so that our system will remember you when you return. They may be set by us or by third party providers whose services we have added to our pages. These cookies enable us to provide enhanced website functionality and personalization as well as make the marketing messages we send to you more relevant to your interests. If you do not allow these cookies, then some or all of these services may not function properly.

Cookies Details

Advertising Cookies

- [x] Advertising Cookies 

These cookies record your visit to our websites, the pages you have visited and the links you have followed to influence the advertisements that you see on other websites. These cookies and the information they collect may be managed by other companies, including our advertising partners, and may be used to build a profile of your interests and show you relevant advertising on other sites. We and our advertising partners will use this information to make our websites and the advertising displayed on it, more relevant to your interests.

Cookies Details

Cookie List

Clear
*   - [x] checkbox label label 

Apply Cancel

Consent Leg.Interest

- [x] checkbox label label

- [x] checkbox label label

- [x] checkbox label label

Decline All Save and Accept

[![Image 3: Powered by Onetrust](https://cdn.cookielaw.org/logos/static/powered_by_logo.svg)](https://www.onetrust.com/solutions/consent-and-preferences/)

Copy debug info
