---
格式版本: 2
标题: "How We Scaled Kimi K2.5 S81695 | GTC San Jose 2026"
原文链接: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/"
发布日期: "2026-08-11"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "Published Time: Tue, 11 Aug 2026 16:37:09 GMT"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 4
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-12T16:09:55+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-12T16:04:00+08:00"
入库时间: "2026-08-12T08:09:55.379Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.nvidia.com/gtc/"
匹配关键词:
  - "GPU"
  - "NVLink"
  - "performance"
  - "throughput"
  - "AI"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 25
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "演讲内容为Kimi模型算法优化与Agent架构，未涉及超节点/AI Rack/机柜级硬件、供电、散热、互连等主题，与项目关注范围无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-12T16:10:05+08:00"
AI主题相关性: 0
AI来源权威性: 12
AI新颖性: 8
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 5
AI摘要: "在GTC San Jose 2026上，演讲者分享了扩展Kimi K2.5模型的方法，重点通过优化token效率、上下文长度和智能体群三个维度提升性能。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:51.452Z"
采集批次: "2026年8月12日13点00分43秒"
采集批次ID: "20260812-130043-786"
去重键: "https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695"
---

Title: How We Scaled Kimi K2.5 S81695 | GTC San Jose 2026

URL Source: https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/

Published Time: Tue, 11 Aug 2026 16:37:09 GMT

Markdown Content:
Visit your regional NVIDIA website for local content, pricing, and where to buy partners specific to your country.

[Continue](https://www.nvidia.com/)

[**GTC Berlin** October 20–22](https://www.nvidia.com/en-eu/gtc/) | [**GTC 2027** March 15–18](https://www.nvidia.com/gtc/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#) 
*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)      
*   [](https://www.nvidia.com/en-us/account/)
*   [Log In](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)[LogOut](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

PLATFORMS

other links

[](https://www.nvidia.com/gtc/)

 Keynote 
*   [Keynote](https://www.nvidia.com/gtc/keynote/)
*   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

 Explore 
*   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
*   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
*   [Speakers](https://www.nvidia.com/gtc/speakers/)
*   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
*   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

[Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)

 More 
*   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
*   [Contact Us](https://www.nvidia.com/gtc/contact/)
*   [FAQ](https://www.nvidia.com/gtc/faq/)
*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*    Keynote 
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*    Explore 
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*    More 
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

*   [](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)

    *   [EN](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
        *   [EN](https://www.nvidia.com/en-us/on-demand/)
        *   [简中](https://www.nvidia.cn/on-demand/)
        *   [日本語](https://www.nvidia.com/ja-jp/on-demand/)
        *   [한국어](https://www.nvidia.com/ko-kr/on-demand/)
        *   [繁中](https://www.nvidia.com/zh-tw/on-demand/)

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/# "Menu")

[Watch On Demand](https://www.nvidia.com/en-us/on-demand/search/?facet.event_name[]=GTC%20San%20Jose&facet.event_year[]=2026&facet.mimetype[]=event%20session&headerText=All%20Sessions&layout=list&page=1&q=-&sort=relevance&sortDir=desc&gtcnavinherit=true)[](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/# "Menu")

*   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [Keynote](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [Keynote](https://www.nvidia.com/gtc/keynote/)
    *   [_GTC Live_ Pregame](https://www.nvidia.com/gtc/pregame/)

*   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [Explore](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [Conference Topics](https://www.nvidia.com/gtc/conference-topics/)
    *   [Poster Gallery](https://www.nvidia.com/gtc/posters/)
    *   [Speakers](https://www.nvidia.com/gtc/speakers/)
    *   [Startups & VCs](https://www.nvidia.com/gtc/startups/)
    *   [Workshops, Training Labs & Certification](https://www.nvidia.com/gtc/training/)

*   [Sponsors & Exhibitors](https://www.nvidia.com/gtc/sponsors/)
*   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [More](https://www.nvidia.com/gtc/session-catalog/sessions/gtc26-s81695/#)
    *   [Code of Conduct](https://www.nvidia.com/gtc/code-of-conduct/)
    *   [Contact Us](https://www.nvidia.com/gtc/contact/)
    *   [FAQ](https://www.nvidia.com/gtc/faq/)
    *   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
    *   [See All GTC Events](https://www.nvidia.com/gtc/events/)

[Video 1](blob:https://www.nvidia.com/87536398-1903-4705-be0e-ca7d260e8732)

Loading

Play

00:00

Play

Seek 10 seconds backwards

Seek 10 seconds forward

00:00 / 00:00

Mute

Press TAB to open volume control Use the arrows to control the volume

Enable Captions

Settings

Play on TV

Turn on Picture in picture

Show Full screen

Transcript Powered by AI

X

X

00:12

Hi everyone, thank you so much for the introduction,

00:15

it's great to be here, it's great to have this opportunity

00:20

to share with you guys some of our latest progress and efforts.

00:26

So one of our major pursuits is to build better open models.

00:34

And we believe in democratizing intelligence.

00:37

With open models, you can deploy anywhere.

00:39

It can be on your local servers.

00:41

It can be on the cloud.

00:42

And you can access every single bit of the weights in the model

00:47

instead of just using a black box.

00:50

And this is one of the slides that I took from Jensen's

00:54

talk earlier this year at CES.

00:56

So as you can see, open models are quickly closing the gap

01:00

with proprietary models and it's reaching the frontier.

01:05

And we believe that with better met open models, we're going

01:07

to make intelligence more accessible to anybody in the world,

01:12

in every corner of the world.

01:15

But open models cannot be just open.

01:18

They have also to be great.

01:22

So in this talk, we're going to discuss how we

01:25

make open models great.

01:28

So as we know, scaling is a primary driver for a lot

01:34

of progress, maybe all of the major AI developments that we have

01:38

witnessed in the last few years.

01:41

And here, we're going to discuss how we scale our model in

01:44

different dimensions.

01:46

So on the left-hand side, the first figure you see here is

01:49

kind of the standard scaling law.

01:52

So on the x-axis, you have the log of the number of training

01:56

tokens, and on the y-axis, you have the log loss.

02:00

And as you scale the number of training tokens, you

02:02

get a lower loss.

02:03

But here the point is, we're not going to just scale the

02:07

number of training tokens, but we also want to improve the token

02:12

efficiency, meaning that we want to move this curve to the left-hand

02:17

side so that we can achieve a lower loss, a much lower loss, using

02:22

the same number of training tokens.

02:25

And this can be achieved by having better architectures

02:28

and optimizers, as we'll discuss in our later slides.

02:33

And the second scaling dimension that we're very interested

02:36

in is to scale the context length.

02:39

So as you can see in the second figure, if we increase

02:43

the context length,

02:45

Then we can have a much higher accuracy in terms of predicting

02:50

the token loss at a given position.

02:55

And this means that we can increase the capability of the

02:58

model to achieve more complex tasks by increasing the context stance.

03:04

So this is the second scaling dimensions that

03:06

we're going to talk about.

03:08

And the third scaling dimension is the number of agents.

03:11

So we introduced this new learning paradigm of agent swarms, where

03:15

we don't just rely on a single agent, but we also orchestrate a

03:21

swarm of agents that can accomplish the subtask in parallel, so that

03:25

we can increase the task capacity.

03:28

And we can translate all of this into the language of agents.

03:32

So if you look at token efficiency, it's mostly about having a

03:36

stronger prior so that you can have more efficiency when you do agent

03:41

RL to search for a better solution.

03:45

And when you think about long context, it's mostly about

03:48

increasing the context so that you can have a longer-running agent.

03:51

It can probably run for days or even weeks or months to accomplish

03:57

more tasks, more complex tasks.

04:01

And for agent swarms, it's another dimension that's added,

04:05

and at the end of the day, we're going to have a swarm of agents.

04:08

That each of them have a super long context, and each of them have a

04:11

very strong prior for us to search in this entire agent RL system.

04:19

All right, so we're going to start from token efficiency.

04:23

So this is one of the most classical figures in the history

04:28

of machine learning.

04:29

So it's taken from Kaplan et al.

04:32

And it basically says that if we scale proportionately

04:35

the number of training tokens, the model parameters, and

04:39

also the amount of compute, we can get lower and lower loss.

04:43

And this is one of the major breakthroughs that

04:47

the entire community.

04:48

has achieved in the last few years to get better intelligence.

04:53

But here, what we're interested in is to have better and better

04:58

token efficiency.

05:00

And here's the thing.

05:02

So one thing that I would like to emphasize is that token efficiency

05:07

is not just about efficiency.

05:09

It's actually also about improving the upper bound of intelligence.

05:14

So here's why.

05:16

So suppose you have, say, 50 trillion tokens, 50 trillion

05:22

high-quality tokens, and then you apply this new optimizer,

05:26

maybe the new optimizer, and then all of a sudden, you

05:29

have a two-time token efficiency.

05:31

So it means that

05:33

It's almost like magic that you get equivalently 100 trillion

05:37

tokens and nowadays we are scaling towards the data war

05:42

and we're hitting the data war and the amount of high quality data

05:47

is quite limited and if we suppose that is a constant amount then

05:51

we increase the token efficiency

05:53

It means that we're going to get better intelligence out of it.

05:57

It's not just about infrastructure efficiency.

05:59

It's about better intelligence.

06:03

So this is why we spend a lot of efforts in this aspect,

06:08

because it's going to push the frontier of intelligence.

06:12

And Mule Optimizer is one of the things that we have heavily

06:15

invested in since last year.

06:19

So it's a second-order optimizer.

06:22

And basically, every single gradient update is transformed

06:27

in a way that each entry is orthogonal to each other.

06:31

And this is very different from the traditional Adam optimizer.

06:35

And if you implement this optimizer properly, you can get a two-time

06:39

token efficiency improvement.

06:42

So we are the first work, we published the first work,

06:48

to demonstrate that our own optimizer is actually scalable

06:52

for LLM training.

06:53

And these are two key techniques that we employ to make it

06:58

effective for large-scale training.

07:00

So one of them is weight decay.

07:02

It is critical for scaling to larger models.

07:05

And the second is we want to ensure a consistent RMS

07:09

update compared to Atom.

07:10

So we have this adjustable coefficient that is applied to each

07:16

update, so that the resulting RMS is going to be comparable to Atom.

07:22

And to make MIOM memory efficient across all these NVIDIA GPU

07:28

clusters, we also developed a distributed MIOM optimizer

07:32

implementation that partitions the states across the data

07:36

parallel group so that we can have a very efficient implementation

07:41

for the MIOM optimizer.

07:43

And these are some of the results that were presented in the paper.

07:47

So as you can see, with the same number of parameters

07:51

and the same number of training tokens, we just replaced the

07:54

original add-on W optimizer with the new Mule optimizer.

07:58

It's going to improve the performance across

08:01

the board significantly.

08:04

But there was this new challenge that we encountered when we

08:09

tried to scale data further, when we tried to scale muon

08:13

for a 1 trillion parameter model, we encountered a new

08:16

issue about training instability.

08:19

So, as you can see on the left figure...

08:22

We observed that the MaxLogix quickly explodes

08:27

and quickly exceeds 1,000.

08:29

And the typical values for training for this MaxLogix is about,

08:36

say, 50 or maybe less than 100.

08:39

But for Mule, it quickly exceeds 1,000.

08:43

And at the same time, we observe training divergence

08:47

on the left-hand side.

08:47

If you look at the training loss, it goes down a bit, but then

08:51

at the end of the day, it explodes, and it cannot converge as expected.

08:56

So this is one of the technical challenges that we have to adjust.

09:01

And the solution to this is to introduce this new

09:04

technique called QKClip.

09:06

So basically what it says is that for each attention head in

09:10

this entire neural network, we're going to, in the forward pass,

09:13

we're going to compute the MaxLogit and then we're going to calculate a

09:18

dividing factor that can be applied to each key projection, as well

09:23

as the query projection, so that we can sort of clip the maximum value

09:30

of the query and the key to sort of Constrainted into a given range.

09:36

So that we're not going to have explosion anymore.

09:40

So these are some of the empirical results.

09:42

On the left-hand side, there are two curves, but they are strictly

09:46

overlapped with each other.

09:47

So these are the training curves before and after applying

09:51

the clipping technique.

09:52

So you can see the clipping technique does not affect

09:56

The training loss decrease at all, but on the right-hand side, if we

10:02

inspect the intermediate metric, if we inspect the MaxLogit, it's going

10:07

to be effectively constrained.

10:10

So it first explodes, as before, but at the value of 100, it's

10:16

going to be clipped at a constant value for a long time, and then

10:19

after a certain number of steps, it will just naturally go down.

10:23

So, the neural networks sort of find a way to constrain the

10:28

maximum value of the MaxLogic to ensure a stable training process.

10:34

And at the same time, it doesn't affect the training convergence

10:38

as shown in the left figure.

10:41

So we employed this technique in our K2 model training and

10:46

successfully scaled it to 1 trillion parameters.

10:50

And this is the first example of large-scale muon training

10:55

in the history of machine learning.

10:58

And the second dimension that we're very interested in is long context.

11:04

So this is another figure.

11:06

It's probably less known.

11:08

It's one of the hidden gems in these papers.

11:12

So instead of just pushing down the training laws by training

11:16

on more tokens, it has some insights from another perspective.

11:21

So as we can see, this is a comparison between

11:24

transformers and LSTMs.

11:27

So on the left-hand side, we can see that transformers

11:29

achieve a lower training loss given the same number of parameters

11:33

and the same number of training tokens as expected.

11:36

And this is why transformers become the de facto architecture

11:41

that people are using right now.

11:42

But on the right-hand side, it's really interesting to

11:45

see that transformers are actually better because it can

11:49

improve through the whole context.

11:52

So the x-axis is the token index in context.

11:54

And if you increase the token index, you can see that the

11:57

training loss of transformers actually dropped by a lot.

12:02

If you just continually increase the context length, the loss

12:06

just continuously drops down.

12:08

But if you look at the curve of LSTM, it just is saturated

12:13

after a certain number of tokens.

12:16

It means that transformers have this better capability

12:19

of capturing longer context, and this is what makes it

12:24

better, because if you go back to like 10 years ago, people used LSTM

12:30

for tasks like machine translation.

12:32

But it is not good for, for example, understanding entire code

12:35

base or running a super long agent trajectory to solve, for example,

12:42

writing Linux kernels from scratch.

12:45

It's not going to be accomplished by LSTMs.

12:47

So this is a very much needed capability in the era of agents,

12:53

because tasks are becoming harder and harder, and we

12:56

need longer and longer contexts.

12:58

So the research idea here is to develop a better architecture

13:03

so that we can efficiently scale to a longer context length and at

13:09

the same time achieve a lower per token loss at larger token indices.

13:16

And this is the motivation for which we introduced this new

13:21

architecture called Kimi linear.

13:24

And it contains this new linear attention variant called Kimi

13:31

delta attention, which improves the original gated delta rule,

13:37

GDR, by improved recurrent memory.

13:40

I will show the details later.

13:41

And at the same time, we're going to mix linear attention

13:44

layers with full attention layers using a 1-2-3 ratio

13:48

so that you can balance between this long context capabilities

13:53

and at the same time having a more efficient implementation.

13:59

So this is some of the formulation.

14:02

The basic idea is simple.

14:03

If you look at linear attention, in the original formulation,

14:08

the memory is going to be global.

14:09

So there is a global single decay factor that

14:14

is applied along the way.

14:16

So it means that

14:18

Basically, there are only two cases.

14:22

In one case, you're going to forget basically everything,

14:25

and you're not going to retain any information.

14:27

In the second case, you can choose to retain almost everything,

14:31

but at the same time, you don't have the capability to

14:34

leave out some of the unnecessary information in this long context.

14:38

So we introduced this key idea of having a fine-grained

14:43

decay factor as shown in this highlighted alpha term.

14:47

So instead of being a scalar, it's going to be a diagonal

14:52

matrix which controls the decay rate for each channel so

14:57

that we can have two possibilities.

14:59

For some of the channels, we can decay really, really

15:02

slow, meaning that we can retain this long-context information

15:06

across a very long range.

15:09

And at the same time, for the other channels, we can

15:12

quickly forget the information from the past indices to refresh

15:17

it and observe new information.

15:20

And this is to increase the expressivity of this model.

15:27

And of course, to leverage modern GPUs, we have to use

15:31

this chunkwise formulations so that we can parallelize

15:36

the computation of modern GPUs.

15:38

So the first equation here is the chunkwise

15:42

formulation of Kimi linear.

15:45

But as you can see, this is going to bring massive infrastructure

15:50

challenges because of this newly introduced alpha term.

15:54

Because now it is a matrix instead of a scalar, it cannot

15:58

easily be factored out.

15:59

So to achieve an efficient implementation, we rewrite

16:04

the entire equation into the bottom three equations.

16:09

So we introduced this matrix inversion operation, as

16:13

well as introducing

16:15

The cumulative decay factor, so that we can implement this

16:19

entire thing in parallel, without sacrificing any efficiency.

16:25

And more importantly, this is not an approximation, it's

16:29

an exact mathematically equivalent formulation, so that we can achieve

16:34

much efficient implementation without sacrificing any loss

16:38

in terms of performance.

16:41

So it's going to be as efficient as previous linear attention

16:47

variants, but at the same time, much more expressive.

16:51

So these are some of the results that we obtained

16:53

using a fair comparison.

16:56

So on the left-hand side, we see the performance on

16:59

two different types of tasks.

17:01

So MMMU is a short context task.

17:04

So for short context tasks, Kimi linear achieved

17:07

a better performance compared to MLA and GDN.

17:12

And at the same time, for longer context tasks such as RULER,

17:16

Kimi linear is also better than the other variants while being much

17:23

more efficient compared to MLA.

17:26

And when we scale the context further to, for example, 1

17:29

million tokens or even longer, it's going to be much more efficient

17:34

compared to the baselines.

17:36

And this is also the first architecture that can outperform

17:41

full attention across the board, including short context tasks, long

17:46

input tasks, and long output tasks.

17:50

So these are two key dimensions that we are interested in.

17:55

And the third dimension is the agent swarms.

17:59

So here is a diagram to showcase how we designed this agent

18:04

swarm paradigm to solve some of the more complex tasks

18:09

compared to single agent paradigms.

18:12

So here we have an orchestrator.

18:15

Or you can call it a main agent.

18:17

It's responsible for orchestrating tasks.

18:20

It has different options.

18:22

For example, we can spawn a group of sub-agents and assign

18:25

new tasks to these sub-agents.

18:28

Or you can collect the results from the return of these sub-agents.

18:33

And you can sort of perform this process in an iterative way.

18:38

And at the end of the day, you can accomplish a more complex task

18:43

compared to using one single agent.

18:45

And it's analogous to human society.

18:48

For example, if we build a company,

18:50

We need different roles and we need, for example, an orchestrator

18:55

or maybe we need a CEO to compose and assign the tasks

18:59

to different roles.

19:00

And at the end of the day, the entire organization

19:03

is going to have to move towards this same goal.

19:07

And here, for example, in this case, we have, maybe

19:10

you have the AI researchers, you have the web developers,

19:14

you have physical researchers, and they can study different topics.

19:17

And at the end of the day, you just collect the results

19:20

and spawn a group of fact-checkers and web developers and file

19:24

downloaders to assemble the results into a single report.

19:32

And this is another perspective to look at this new paradigm.

19:38

So the x-axis is the complexity of the task.

19:44

And the y-axis is the execution time.

19:46

And the complexity of the task is measured by the accuracy

19:51

of a group of models on such tasks.

19:55

So we can see with Agent Swarms, it's going to substantially

20:01

reduce the execution time compared to single agents.

20:08

It's going to be more effective, and this means

20:11

that we can scale this Agent Swarm paradigm Time to...

20:15

For example, if you run Agent Swarms with 100 or maybe even 1,000

20:20

sub-agents, you can accomplish a complex task within a certain

20:25

period of time that is tolerable for producing real economic value.

20:33

And we can certainly scale it in different dimensions.

20:36

We can scale the inputs.

20:38

For example, we can download and read hundreds of sources

20:42

or even maybe thousands of doses in parallel.

20:45

Or we can output, write a 100-page literature review in parallel.

20:52

Or we can take actions at scale.

20:54

We can perform data analysis for 10 different tasks.

20:58

And also it is orchestration at scale.

21:01

You have to learn to design subtasks and aggregate the results.

21:06

And technically, we define some new objective functions

21:10

to guide the learning process of our agent swarm system.

21:15

So there are three reward functions, reward objectives

21:20

that are considered here, compared to the conventional

21:25

single agent IRL learning.

21:27

So the first term is what we call the instantiation reward.

21:32

It incentivizes sub-agent instantiation to prevent

21:37

this serial collapse phenomenon from happening.

21:42

So basically, we don't want it to default to single-agent execution.

21:47

We want to encourage the parallel executions, especially when we...

21:54

When it's early stage in training and of course we can decay

21:59

the weight for this instantiation reward term over training course

22:04

because when it learns parallel execution we can reduce the weight.

22:10

And the second term here is Finish Reward.

22:14

And it is used because we observe one of the things in training

22:19

that some of these subtasks are just created but never finished.

22:25

So it's almost like it's going to hack the first term.

22:28

By just spawning a bunch of sub-agents, and the task might

22:31

be too complex, or maybe the task just doesn't make sense.

22:35

And here, we use this Finish Reward to basically encourage

22:39

that each of the sub-tasks should have a relatively high

22:44

ratio of completion.

22:46

Instead of just spawning a bunch of pseudo-tasks, we

22:50

need it to be meaningful.

22:52

So this is the second term that we use and of course we use

22:55

the same, you know, decay strategy.

22:58

We use the relative high weight at the beginning of training

23:01

and we decay it to a relatively low weight at the end of training.

23:05

And of course the third term is the standard term.

23:07

It's the outcome reward.

23:09

It's going to measure whether the entire task is completed and

23:15

then we're going to add these three terms in our reinforcement learning

23:21

And of course, we have to build the entire infrastructure

23:24

because right now you need to support the parallel execution

23:28

and they will need to support different reward functions

23:31

and to maximize the efficiency of the entire agent swarm IO system.

23:38

So here are three different things that we have tried scaling.

23:45

The Muon Clip Optimizer improves token efficiency.

23:49

And Kimi Delta attention in the Kimi linear architecture

23:53

improves long contacts.

23:55

And we also have the agent swarms paradigm to further

24:00

create a new dimension of scaling.

24:03

And all of this put together, we created Kimi K2.5, a new

24:08

model that we just released over one month ago.

24:12

Here's a short video to demonstrate some of its capabilities.

25:19

So yeah, there are a lot of interesting things, capabilities

25:23

that we discover from the model.

25:25

For example, it merges the visual capabilities with

25:29

coding capabilities.

25:30

So a lot of new things just emerge out of it.

25:34

It can read a video and then produce a website that sort

25:39

of replicates or style transfer The original video.

25:43

And all of this are due to successful and stable training

25:49

at the pre-training stage.

25:50

So this is also one of the most beautiful curves that

25:55

I observed in my life.

25:56

So this is the training curve of the K2.5 based model.

26:02

So as you can see, it went through over 15 trillion tokens and

26:05

of course in K2.5 we additionally trained another 15 trillion

26:09

tokens and the entire training process is just so stable,

26:14

there's no lost spike, especially when we introduced this new

26:18

Muon optimizer, we didn't observe any spike and this

26:21

smooth, stable training process produces a very stable outcome.

26:27

A very strong-based model that we can fine-tune on top of

26:30

it to achieve new capabilities as we introduced and saw in the video.

26:37

And this is also, of course, a trend on NVIDIA H800 GPUs, and you

26:42

should know in this H800 cluster contains two TB RAM and 8 GPUs.

26:50

They are connected by NVLink.

26:52

And another key innovation of KeyMeK 2.5 is that it is

26:59

the first open model with native joint vision text capabilities.

27:04

So if you look at previous open models, usually their

27:08

visual capabilities are added on top of a text base, meaning

27:12

that, for example, if you train the text models for

27:14

20 trillion tokens, and then on top of it, you do another 2 trillion.

27:19

Sort of a post-training process to add additional

27:23

visual capabilities on top of it.

27:25

But for K2.5, it's different in the sense that we fuse

27:30

the training process of vision and text from day one.

27:34

So it's called early fusion here.

27:36

We start from 0% of the progress.

27:39

So from day one, we're going to merge the vision and text

27:43

tokens and as shown in our preliminary experiments, it

27:47

outperforms our late fusion.

27:50

And some of the new capabilities that we observe also come

27:53

from this training recipe.

27:55

For example, if you want to do...

27:57

Vision to kill, you really have to merge vision and text into

28:02

a single brand to achieve that.

28:04

If you separate these two brands, it's not going to happen.

28:07

You have to align these two modalities into a shared embedding

28:11

space, a shared representation space so as to achieve this.

28:18

And another interesting thing that we observe is that these

28:22

two modalities can actually enhance each other.

28:25

So that's been long been a challenge that if you add

28:30

vision capabilities into a text model, it's going to

28:33

somewhat hurt the text performance.

28:37

But here we found that if you train it properly, these two modalities

28:41

can actually enhance each other.

28:43

So this is one of the key findings that

28:45

We observe in our training, so first, vision improves text.

28:50

So this is so interesting, so before Vision RL,

28:54

The performance in the first column, and then we have the

28:57

performance after Vision RL.

28:59

So here, Vision RL refers to a process that we only

29:04

use vision tasks.

29:05

So there is no task involved here.

29:09

We only have vision tasks.

29:10

For example, we teach the model how to count, how to answer

29:14

some of these visual QA problems without any, for example, math and

29:20

encoding problems in this space.

29:22

But we observe that it's going to improve the performance

29:25

for even, you know, reasonably heavy text tasks.

29:29

And on the other hand, text also improves vision.

29:32

If you have a very strong text space, you actually don't

29:35

need any vision SFT data in the training process.

29:39

And this is the approach that we adopt.

29:41

So it's called Zero Vision SFT.

29:44

Basically, we don't have...

29:45

We have basically zero vision SFT data, and the only SFT data

29:49

that we have is the text SFT data.

29:52

And then we do a joint RL over text and vision, and

29:56

you can see that we can achieve almost state-of-the-art performance

29:59

across the board on vision tasks without any vision data.

30:03

So it's clear that if you have a strong text space, it's also

30:07

going to improve the vision if you align these two modalities into a

30:13

shared space in your pre-training.

30:17

And also, these are some of the examples of

30:26

As I've shown in the video, it demonstrates strong capabilities

30:30

of visual design and front-end coding, and this also emerges from

30:35

our vision text through training.

30:39

So after all this, so this is all about Kimi K2.5.

30:44

And as you probably know, we released our new architecture

30:50

yesterday in our tech report.

30:52

It's called Attention Residue.

30:54

So here I'm also going to briefly talk about our new

30:58

work, which serves as a sneak peek into our next generation

31:02

architecture that we're probably going to adopt in our later models.

31:07

So here the motivation is quite simple.

31:11

Can we apply some of our techniques that we use in the temporal

31:16

dimension, and we just take some of the inspirations and then

31:21

apply it to the depth dimension?

31:26

And it starts from this residual connection.

31:28

So I still remember listening to Kimi's talk.

31:34

at a tutorial in ICML 2016, 10 years ago.

31:39

So it was a brilliant idea.

31:40

So basically, before ResNet, nobody was able to train deep networks.

31:47

If you increase the depth, if you increase the number

31:49

of layers for neural networks,

31:52

Nobody was able to train it because you observe this gradient

31:55

exposure and gradient vanishing, all these stability issues.

31:59

But then after the introduction of ResNet, we can train an

32:04

arbitrarily large number of layers.

32:06

You can stack as many layers as you want, and you don't

32:10

have to worry about the training stability issue and stuff.

32:14

And as discussed in Ilya's talk two years ago,

32:18

It basically says that residual connection is a variant of

32:23

LSTM, but just rotated 90 degrees.

32:27

So how do you understand this?

32:28

If you look at LSTM, it's a variant of recurrent net, right?

32:33

And it's a recurrent model process.

32:36

So we're going to take the hidden states from the last step,

32:41

and then we're going to have some gating mechanism, some function

32:45

to produce the current states.

32:48

And if you look at the depth dimension, the register connection

32:52

is basically the same.

32:53

We're going to take the outputs from the last layer, and then

32:56

we're going to apply some sort of function on top of

32:59

it to produce the current outputs of the current layer.

33:06

It's just the formulation is different.

33:08

For example, for a residual connection, we're going to

33:10

use a fixed addition.

33:11

We're going to add the previous hidden state

33:18

with the current output.

33:20

It's just the formulation that's different.

33:22

But the basic idea is the same.

33:23

It's a recurrent that applies in the dimension of death.

33:28

But on the other hand...

33:30

We can think about reformulating this function, instead of

33:36

having an LSTM, can we have an attention in the dimension

33:40

of that, and it's going to create new possibilities because attention

33:45

has been demonstrated to be so successful in the transformer era.

33:50

So what we're going to do is not just to take the last

33:53

hidden state, but we're going to consider all the previous

33:56

hidden states and use the attention operation, the attention mechanism,

34:00

to assemble and aggregate all of these previous hidden states

34:05

to compute the current state.

34:07

So this is exactly attention rotated by 90 degrees.

34:13

We view it as a natural generalization of residual

34:17

connections in the LSTM analogy.

34:21

Okay, and here is the detailed formulation.

34:24

So, on the left-hand side is a standard residue.

34:27

As I said, it is basically LSTM rotated by 90 degrees.

34:31

And the second figure is attention rotated by 90 degrees.

34:35

So, what we did is to collect all the previous hidden states

34:39

and have a simple attention operation on top of it to

34:43

produce the current layers outcome.

34:45

And of course, to increase the efficiency, to reduce

34:49

the infrastructure, for example, communication and memory overhead,

34:53

we also designed a new variant called block attention residue

34:57

on the right-hand side.

34:59

So basically, the idea is also simple.

35:02

We're going to divide all the layers in the neural networks

35:06

into multiple blocks.

35:08

For example, each block can contain, say, 16 layers, or

35:11

it can contain maybe four layers.

35:13

And then for each block, we're going to apply this

35:18

attention residue only on the output of each block.

35:21

But within each block, we also still adopt this standard residue.

35:25

So this is going to reduce a lot of overhead while having minimal

35:30

loss in terms of training accuracy.

35:33

And these are some of the impressive results that we

35:35

achieved on this new architecture.

35:39

So, on the scaling law, we can improve the token efficiency

35:44

by 24%, meaning that if you have 50 trillion high-quality

35:50

tokens, now you just magically have over 60 trillion tokens.

35:58

For the validation loss, you can also observe that it's

36:02

consistently lower than the original curve, demonstrating this

36:08

stability across optimisation, and also achieves the best improvement

36:14

on some of these coding, math, and reasoning heavy tasks,

36:18

as shown in the benchmark results of GPQA, Math, and HumanEval.

36:25

So the entire community keeps moving forward and we're happy that

36:30

we can, we're able to contribute to the community with new technologies

36:35

and some of these technologies have been sort of standard

36:41

and de facto for a long time, but as you can see, we still see a

36:45

lot of opportunities to improve it, to have revolutionary new design

36:52

to achieve better performance.

36:53

If you multiply all these scans together, you can actually

36:57

have a much better model.

36:59

So Adam was invented in 2014, and now we scale an

37:04

open-source MuonClip, a drop-in replacement for Adam.

37:09

And I'm sure that if you're training Transformer LLM,

37:14

it's going to be much better if you use MuonClip instead of Adam.

37:18

And Attention was invented over eight years ago.

37:22

And then now we have Kimi linear, which is a linear version.

37:25

We don't have to use full attention across all layers.

37:29

We can have linear attention that performs better on short,

37:32

long contacts at the same time.

37:35

And also, residual connections are now

37:39

Also a challenge, we scaled an open-source attention residue.

37:44

So I think one of the interesting things about our error is

37:49

that we sort of adopt a different mindset for doing research.

37:54

So if we go back to 10 years ago, it's mostly about publishing a

37:58

new idea, but then I think the lack of the rigor of the experiments

38:04

It's very hard to produce solid experimental results.

38:09

But now we have this scaling ladder.

38:10

We have enough resources to train the model and run it

38:16

at different scales.

38:17

We can have a whole set of benchmarks to measure the progress.

38:21

So it becomes easier to make confident and solid

38:26

conclusion out of it.

38:27

And this is one of the reasons why we are observing New

38:31

progress on this.

38:34

Ancient techniques, and I'm sure that we'll see more and

38:36

more, especially in the open source community.

38:40

I think we're going to have more and more even better

38:43

architectural and optimization improvement in the next few years.

38:50

All right, so to summarize, we're going to keep scaling our models.

38:54

And so these are three dimensions.

38:58

For example, we see different architectures and optimizers

39:04

that optimize all three dimensions.

39:07

And we'll keep. We see new dimensions for scaling, agent

39:12

swarms is not the end, and we are glad that we can move

39:16

forward with the entire open source community to achieve

39:20

better and better intelligence.

39:23

Thank you so much.

# How We Scaled Kimi K2.5

Zhilin Yang,Founder and CEO,Kimi (Moonshot AI)

AI-Generated Summary of this Video

Rate Now

Explore how we scaled Kimi K2.5 by pioneering the Muon optimizer to double token learning efficiency and maximize training throughput through Day 0 infrastructure co-design. Gain insights into AI-native training and linear attention architectures to help you unlock the potential of longer-running agents.

### Learn More About This Topic

Share

Favorite

Add to list

PDF 

Events & Trainings:GTC San Jose

Date:March 2026

Industry:All Industries

Topic:Developer Tools & Techniques - Large Language Models (LLMs)

Level:General Interest

Language:English

NVIDIA technology:Blackwell, Hopper

Region:

Company Information

*   [About Us](https://www.nvidia.com/en-us/about-nvidia/)
*   [Investors](https://investor.nvidia.com/home/default.aspx)
*   [Venture Capital (NVentures)](https://www.nvidia.com/en-us/startups/nventures/)
*   [NVIDIA Foundation](https://www.nvidia.com/en-us/foundation/)
*   [Research](https://www.nvidia.com/en-us/research/)
*   [Corporate Sustainability](https://www.nvidia.com/en-us/sustainability/)
*   [Technologies](https://www.nvidia.com/en-us/technologies/)
*   [Careers](https://www.nvidia.com/en-us/about-nvidia/careers/)

News and Events

*   [Newsroom](https://nvidianews.nvidia.com/)
*   [Company Blog](https://blogs.nvidia.com/)
*   [Technical Blog](https://developer.nvidia.com/blog/)
*   [Webinars](https://www.nvidia.com/en-us/about-nvidia/webinar-portal/)
*   [Stay Informed](https://www.nvidia.com/en-us/preferences/email-signup/)
*   [Events Calendar](https://www.nvidia.com/en-us/events/)
*   [GTC AI Conference](https://www.nvidia.com/gtc/events/)
*   [NVIDIA On-Demand](https://www.nvidia.com/en-us/on-demand/)

Popular Links

*   [Developers](https://developer.nvidia.com/)
*   [Partners](https://www.nvidia.com/en-us/about-nvidia/partners/)
*   [Executive Insights](https://www.nvidia.com/en-us/executive-insights/)
*   [Startups and VCs](https://www.nvidia.com/en-us/startups/)
*   [Documentation](https://docs.nvidia.com/)
*   [Technical Training](https://www.nvidia.com/en-us/learn/organizations/)
*   [Professional Services for Data Science](https://www.nvidia.com/en-us/support/enterprise/advisory-services/)

Follow NVIDIA 

[](https://www.facebook.com/NVIDIA "<util:I18n key=\"Follow GeForce on Facebook\" />")[](https://www.instagram.com/nvidia/?hl=en)[](https://www.linkedin.com/company/nvidia/)[](https://twitter.com/nvidia "<util:I18n key=\"Follow GeForce on Twitter\" />")[](https://www.youtube.com/user/nvidia)

[United States](https://www.nvidia.com/en-us/location-selector/)

*   [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/)
*   [Your Privacy Choices](https://www.nvidia.com/en-us/about-nvidia/privacy-center/)
*   [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/)
*   [Accessibility](https://www.nvidia.com/en-us/about-nvidia/accessibility/)
*   [Corporate Policies](https://www.nvidia.com/en-us/about-nvidia/company-policies/)
*   [Product Security](https://www.nvidia.com/en-us/product-security/)
*   [Contact](https://www.nvidia.com/en-us/contact/)

Copyright © 2026 NVIDIA Corporation

Select Location

The Americas

*   [Argentina](https://www.nvidia.com/es-la/ "Argentina")
*   [Brasil (Brazil)](https://www.nvidia.com/pt-br/ "Brasil (Brazil)")
*   [Canada](https://www.nvidia.com/en-us/ "Canada")
*   [Chile](https://www.nvidia.com/es-la/ "Chile")
*   [Colombia](https://www.nvidia.com/es-la/ "Colombia")
*   [México (Mexico)](https://www.nvidia.com/es-la/ "México (Mexico)")
*   [Peru](https://www.nvidia.com/es-la/ "Peru")
*   [United States](https://www.nvidia.com/en-us/ "United States")

Europe

*   [België (Belgium)](https://www.nvidia.com/nl-nl/ "België (Belgium)")
*   [Belgique (Belgium)](https://www.nvidia.com/fr-be/ "Belgique (Belgium)")
*   [Česká Republika (Czech Republic)](https://www.nvidia.com/cs-cz/ "Česká Republika (Czech Republic)")
*   [Danmark (Denmark)](https://www.nvidia.com/da-dk/ "Danmark (Denmark)")
*   [Deutschland (Germany)](https://www.nvidia.com/de-de/ "Deutschland (Germany)")
*   [España (Spain)](https://www.nvidia.com/es-es/ "España (Spain)")
*   [France](https://www.nvidia.com/fr-fr/ "France")
*   [Italia (Italy)](https://www.nvidia.com/it-it/ "Italia (Italy)")
*   [Nederland (Netherlands)](https://www.nvidia.com/nl-nl/ "Nederland (Netherlands)")
*   [Norge (Norway)](https://www.nvidia.com/nb-no/ "Norge (Norway)")
*   [Österreich (Austria)](https://www.nvidia.com/de-at/ "Österreich (Austria)")
*   [Polska (Poland)](https://www.nvidia.com/pl-pl/ "Polska (Poland)")
*   [România (Romania)](https://www.nvidia.com/ro-ro/ "România (Romania)")
*   [Suomi (Finland)](https://www.nvidia.com/fi-fi/ "Suomi (Finland)")
*   [Sverige (Sweden)](https://www.nvidia.com/sv-se/ "Sverige (Sweden)")
*   [Türkiye (Turkey)](https://www.nvidia.com/tr-tr/ "Türkiye (Turkey)")
*   [United Kingdom](https://www.nvidia.com/en-gb/ "United Kingdom")
*   [Rest of Europe](https://www.nvidia.com/en-eu/ "Rest of Europe")

Asia

*   [Australia](https://www.nvidia.com/en-au/ "Australia")
*   [中国大陆 (Mainland China)](https://www.nvidia.com/zh-cn/ "中国大陆 (Mainland China)")
*   [India](https://www.nvidia.com/en-in/ "India")
*   [日本 (Japan)](https://www.nvidia.com/ja-jp/ "日本 (Japan)")
*   [대한민국 (South Korea)](https://www.nvidia.com/ko-kr/ "대한민국 (South Korea)")
*   [Singapore](https://www.nvidia.com/en-sg/ "Singapore")
*   [台灣 (Taiwan)](https://www.nvidia.com/zh-tw/ "台灣 (Taiwan)")

Middle East

*   [Middle East](https://www.nvidia.com/en-me/ "Middle East")

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/). You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control (GPC) signal and recorded your rejection of all optional cookies on this site for this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. To opt out of non-cookie personal information "sales" / "sharing" for targeted advertising purposes, please visit the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information on our privacy practices.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies. You can manage these settings in the NVIDIA [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

We have detected the Global Privacy Control Signal (GPC) and have opted you out of all optional cookies on this browser. You can manage your cookie settings by clicking on "Manage Settings". Please see our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) for more information. We have also opted you out of "sharing"/"sales" of personal information outside of cookies which overrides at least one of your previous settings. You can manage them in the [NVIDIA Preference Center](https://www.nvidia.com/en-us/privacy-center/). Please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) for more information.

Manage Settings

Reject Optional Accept All

![Image 2: Company Logo](https://cdn.cookielaw.org/logos/10ddf4ca-c072-45d0-b3ac-eead0ed93db0/6e17f6e4-c77b-4a11-9f34-c107c42e4bfc/7981.png)

Cookie Settings

We and our third-party partners (including social media, advertising, and analytics partners) use cookies and other tracking technologies to collect, store, monitor, and process certain information about you when you visit our website. The information collected might relate to you, your preferences, or your device. We use that information to make the site work, analyze performance and traffic on our website, provide a more personalized web experience, and assist in our marketing efforts.

Under certain privacy laws, you have the right to direct us not to "sell" or "share" your personal information for targeted advertising. To opt-out of the "sale" and "sharing" of personal information through cookies, you must opt-out of optional cookies using the toggles below. To opt out of the "sale" and "sharing" of data collected by other means (e.g., online forms) you must also update your data sharing preferences through the [NVIDIA Preference Center](https://www.nvidia.com/en-us/about-nvidia/privacy-center/).

Click on the different category headings below to find out more and change the settings according to your preference. You cannot opt out of Required Cookies as they are deployed to ensure the proper functioning of our website (such as prompting the cookie banner and remembering your settings, etc.). By clicking "Save and Accept" or "Decline All" at the bottom, you consent to the use of cookies and other tools as described in our [Cookie Policy](https://www.nvidia.com/en-us/about-nvidia/cookie-policy/) in accordance with your settings and accept our [Terms of Service](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) (which contains important waivers). For more information about our privacy practices, please see our [Privacy Policy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/).

Required Cookies

Always Active

These cookies enable core functionality such as security, network management, and accessibility. These cookies are required for the site to function and cannot be turned off.

Cookies Details

Performance Cookies

- [x] Performance Cookies 

These cookies are used to provide quantitative measures of our website visitors, such as the number of times you visit, time on page, your mouse movements, scrolling, clicks and keystroke activity on the websites; other browsing, search, or product research behavior; and what brought you to our site. These cookies may store a unique ID so that our system will remember you when you return. Information collected with these cookies is used to measure and find ways to improve website performance.

Cookies Details

Personalization Cookies

- [x] Personalization Cookies 

These cookies collect data about how you have interacted with our website to help us improve your web experience, such as which pages you have visited. These cookies may store a unique ID so that our system will remember you when you return. They may be set by us or by third party providers whose services we have added to our pages. These cookies enable us to provide enhanced website functionality and personalization as well as make the marketing messages we send to you more relevant to your interests. If you do not allow these cookies, then some or all of these services may not function properly.

Cookies Details

Advertising Cookies

- [x] Advertising Cookies 

These cookies record your visit to our websites, the pages you have visited and the links you have followed to influence the advertisements that you see on other websites. These cookies and the information they collect may be managed by other companies, including our advertising partners, and may be used to build a profile of your interests and show you relevant advertising on other sites. We and our advertising partners will use this information to make our websites and the advertising displayed on it, more relevant to your interests.

Cookies Details

Cookie List

Clear
*   - [x] checkbox label label 

Apply Cancel

Consent Leg.Interest

- [x] checkbox label label

- [x] checkbox label label

- [x] checkbox label label

Decline All Save and Accept

[![Image 3: Powered by Onetrust](https://cdn.cookielaw.org/logos/static/powered_by_logo.svg)](https://www.onetrust.com/solutions/consent-and-preferences/)

Copy debug info
