---
格式版本: 2
标题: "Next '26: The Future of AI Infrastructure"
原文链接: "https://www.youtube.com/watch?v=PJQPMv8TqLA"
发布日期: "2026-04-22"
发现时间: "2026-04-22T00:00:00+08:00"
入库时间: "2026-05-28T02:39:38+08:00"
来源平台: "历史资料迁移"
搜索渠道: "legacy_migration"
搜索词: "url-index.json 迁移"
匹配关键词:
  - "AI Rack"
  - "Scale-up"
  - "高速互联子系统"
  - "url-index.json 迁移"
  - "待定数据"
  - "厂家"
  - "Google"
  - "YouTube"
  - "Scale"
  - "up"
  - "22.md"
相关厂家:
  - "Google"
相关专家:
  []
内容类型: "视频"
抓取工具: "历史采集"
清洗工具: "历史资料原文保留 + LLM 正文裁剪"
关联判断: "待人工确认"
关联置信度: 0
关联理由: "历史资料迁移，未重新调用模型判定"
AI优质: "是"
AI打分: 88
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "本文系Google高管官方演讲实录，权威性强。详细披露第八代TPU的Pod级架构、万卡规模、自定义ICI互连、Scale-up/Scale-out带宽及97%系统Goodput等核心指标，明确量产部署计划与客户落地案例，高度契合机柜级AI…"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-06-03T00:30:39+08:00"
AI主题相关性: 16
AI来源权威性: 15
AI新颖性: 18
AI技术细节: 16
AI商业部署信号: 14
AI完整性: 9
采集批次: "2026年5月28日2点39分33秒"
采集批次ID: "legacy-migration-20260528023933"
去重键: "https://www.youtube.com/watch?v=PJQPMv8TqLA"
---

## 视频描述

Amin Vahdat，Google高级副总裁、首席技术官（AI与基础设施），深入介绍支持 agentic 时代的基础设施。Acquired 播客的 Ben Gilbert 和 David Rosenthal 将现场采访他，探讨 Google 如何构建未来 AI 创新的架构。

#googlecloudnext

## 字幕全文

**[00:00:10]** 哇，见到大家太棒了。我们今晚真的非常激动，能邀请大家来到这里，这是我们通往 Google Next 接下来几天的重要时刻。基础设施

**[00:00:19]** 是我们所做一切的基础。今天我们会大量讨论 AI，也会谈到智能体（agents）。归根结底，都是基础设施驱动这一切。

![关键帧 00:00:19](./assets/20260424_PJQPMv8TqLA_19s.jpg)

> 视觉分析：Amin Vahdat，SVP 和首席技术官，负责 Google 的 AI 和基础设施。

**[00:00:27]** 提醒一下，今晚你们看到的内容是别人还未看到的。所以请保密到明天。好的，那我想先

**[00:00:36]** 从 Google 的使命谈起。我在 Google 已经有 16 年了，

**[00:00:42]** 还有一个月就满。这个使命一直让我非常激励——组织世界信息，让其普遍可访问且有用。

**[00:00:52]** 我认为这是一个永恒的使命，而真正实现这个使命需要建构之前未曾有过的

**[00:00:58]** 基础设施。25 年前，人们建设基础设施的方式

**[00:01:03]** 我们暂且不做历史回顾，但25年前的基础设施并不支持这个使命，也无法支持网页搜索。创始人和

**[00:01:13]** 公司最早的员工，Larry、Sergey，以及其他人，他们意识到要解决这个问题，

**[00:01:19]** 需要在基础设施和其基本原理上大量投入，才能解决

**[00:01:24]** 网页搜索。如今，过去的 25 年后，实现这个使命需要突破智能本身。

**[00:01:35]** 所以挑战比以往更大，机会比以往更大。就像25年前一样，要解决智能问题的基础设施至今还未

**[00:01:45]** 存在。我们正在构建它。我认为我们取得了非常、

**[00:01:53]** 非常重大的突破。今天你们会了解到，我们将迈出

**[00:01:58]** 一个朝着定义未来基础设施的重要步伐，

**[00:02:06]** 这个智能体时代对基础设施提出了前所未有的需求。这

**[00:02:11]** 不是 AI 生成的图片。这是田纳西州克拉克斯维尔，我们在那里的数据中心。

![关键帧 00:02:11](./assets/20260424_PJQPMv8TqLA_131s.jpg)

> 视觉分析：关键标题：“The agentic era is placing unprecedented demand on infrastructure”，图片显示数据中心。

**[00:02:17]** 这就是支撑我们今天服务的基础。搜索、广告，

**[00:02:24]** YouTube，还有十多个日活超五亿的服务，都在这套

**[00:02:30]** 基础设施上运行。实际上，每一个服务都已经嵌入了 AI。赋能所有

**[00:02:35]** 的当然是我们的自研基础设施。今天我们会重点讨论它。

**[00:02:43]** 这套 AI 堆栈是过去几年专为解决生成式 AI

![关键帧 00:02:43](./assets/20260424_PJQPMv8TqLA_163s.jpg)

> 视觉分析：幻灯片标题：“AI Stack”，提到一个图标表示 AI 技术栈。

**[00:02:56]** 最大挑战而定制开发的。在 Google，我们有独特的机会能自上而下纵向

**[00:03:03]** 集成，真正端到端地解决问题。我们不仅仅插入某一层，

**[00:03:08]** 我们能把整个问题解决。起点是能源。最大约束之一。我们对此非常重视。

**[00:03:16]** 责任和挑战都非常严峻。过去25年甚至更久，我们一直努力扩展以满足全球需求，如今更是巨大的挑战和机遇，

**[00:03:27]** 我们正在进行基础性工作，扩展并负责任地扩展基础设施。

**[00:03:33]** 接下来是数据中心，土地、机房、建筑——你

**[00:03:39]** 看到的冷却系统、机械设备都承载着我们的基础设施，

**[00:03:47]** 还有机架、硬件、服务器、存储、网络、TPU、GPU等，运行着下一层

**[00:03:56]** 是软件，这才让一切协同

**[00:04:02]** 没有软件，硬件其实很难发挥作用。从

**[00:04:08]** 这里开始，过去几年改变最大的是基础模型。Gemini 3，

**[00:04:14]** 业界领先，全部运行在我们自研、自建基础设施上，然后

**[00:04:21]** 这些模型进一步驱动我们的所有服务——搜索、广告、YouTube、照片、Gmail 等等。

**[00:04:29]** 它们还驱动云服务。成千上万的企业依赖 Gemini 进行日常工作。

**[00:04:41]** 我们把底部四层称为基础设施堆栈。事实上，

**[00:04:48]** 如果你拿六层中的任何一层单独设计，最终

**[00:04:54]** 你或许只能达到“最低共同标准”。而我们能做到

**[00:05:01]** 不让任何细节在衔接处流失。从一层到另一层，

**[00:05:10]** 我们确保端到端集成，最大效率、

**[00:05:17]** 最大可靠性和最高安全性。

**[00:05:28]** 现在我们聚焦硬件。今天要重点讲的一部分。我们做驱动 AI 基础设施的硬件已有十多年。2013年启动了

![关键帧 00:05:28](./assets/20260424_PJQPMv8TqLA_328s.jpg)

> 视觉分析：幻灯片标题：“TPU Supercomputing”，展示 TPU 年份演进图（2015-2025）。

**[00:05:36]** 自研 AI 定制芯片——张量处理单元 TPUs。我们

**[00:05:42]** 意识到，要满足未来服务需求，需要一种

**[00:05:49]** 全新的硬件架构，因此我们

**[00:05:56]** 发明了张量处理单元。每一代都不断创新：2015年推出 TPUv1，

**[00:06:04]** 2018年 TPUv2，开发新功能，预示了业界今日

**[00:06:14]** 加速机器学习计算的发展方向，如液冷、自定义数学运算、自定义网络，连接成百上千芯片形成

**[00:06:20]** 超级计算机。创新速度加快，从V1到V2间隔三年，到后来一到两年，甚至每年一代。去年我们在Next发布了Ironwood，

**[00:06:27]** 迄今最强的TPU超级计算机。那么问题来了，

**[00:06:41]** 2026年我们如何继续？Ironwood真正改变了世界。

**[00:06:47]** 现在它在全球大规模应用部署，Google 正在使用，

**[00:06:55]** DeepMind 正在使用，Google Cloud 的众多服务都在运行 Ironwood。

**[00:07:02]** 对2026年，我非常自豪地介绍我们第八代 TPUs。

![关键帧 00:07:02](./assets/20260424_PJQPMv8TqLA_422s.jpg)

> 视觉分析：无实际意义内容。

**[00:07:16]** Google 的命名很直白，你会注意到这里

**[00:07:28]** 只有一个 S。我们并不是发布第八代 TPU，而是发布第八代 TPUs。

**[00:07:35]** 首次，两款自研 TPUs：8T 和 8i，从零开始

**[00:07:43]** 面向时代需求设计。TPU8 T 用于训练，是强大的动力引擎、最大规模。TPU8

**[00:07:52]** 专为推理打造，服务智能体时代，最低延迟，把内容最快送达全球用户。

**[00:08:05]** 下面介绍细节。TPU8，每个 pod 芯片数量差不多，

**[00:08:12]** 但 9,600 个芯片通过自定义 ICI 互连。

**[00:08:20]** 每个 TPU pod 浮点算力几乎是上一代的三倍，

**[00:08:25]** 网络带宽翻倍，Scale-out 可达四倍带宽，

**[00:08:32]** Scale-up 网络带宽也四倍。TPU8i作为推理引擎，pod规模扩大至1152片以上，

**[00:08:41]** 能实现最大模型协同，驱动最强智能体，实时协同，

**[00:08:47]** pod 浮点算力达10倍 exaflop。

**[00:08:54]** 高性能内存容量扩大7倍，scale-up下仍保持。什么

**[00:09:10]** 8倍、10倍全部一年内完成，进步速度令人震惊。

**[00:09:16]** 我们认为，每一年都开启了芯片能力上的全新应用场景。

**[00:09:26]** 所有工程师，包括你们世界各地的工程师，都会

**[00:09:31]** 思考以前看似不可能的问题，现在可挑战。我们开启了

**[00:09:36]** 全新能力、场景和服务世界。无论是搜索、广告、YouTube，

**[00:09:43]** 还是你最喜欢的服务，一年后拥有更多算力能做什么？

**[00:09:53]** 在第八代，值得关注的另一个点是，两款芯片都是

**[00:10:00]** 独立设计，分别用于训练和推理，而非简单

**[00:10:08]** 相互衍生。规格、能力、互连完全不同，针对各自需求。需要训练动力就用8T，需要极速推理就用8i。

**[00:10:22]** 这些 TPUs 将持续驱动 Gemini，但它的意义不限于模型层面。

**[00:10:28]** 模型当然重要，但我们看到的是，通过 Gemini Enterprise 等服务，

**[00:10:33]** 企业运作方式被彻底改变。医疗、科学突破也不例外。

**[00:10:39]** 涨速惊人。换句话说，

**[00:10:45]** 十年的研究现在可能一年完成。

**[00:10:50]** 所能抵达的进步程度主要取决于底层算力。

**[00:11:00]** 很高兴今天与大家分享第八代 TPUs。这两年多

**[00:11:05]** 打磨而成的产品，今年晚些时候会正式上线。也想借此机会

**[00:11:14]** 感谢数千名Google的工程师，他们在这两年为两款芯片努力付出。

**[00:11:21]** 团队辛苦完成了两款芯片。特别让我欣慰的是，几年前就意识到一年一款芯片已不够，

**[00:11:26]** 这是首次挑战两款高性能定制芯片，团队也成功交付。

**[00:11:32]** 所以在此，请

**[00:11:37]** Ben 和 David 回到舞台，我们将一起讨论这些芯片。谢谢大家！

**[00:11:43]** 欢迎回来！很高兴见到你。谢谢。两枚芯片，两枚芯片。两枚芯片。一

**[00:11:51]** 年，双倍乐趣。你确实很忙。真是一年忙，同时也是一年充满乐趣。我们知道

**[00:11:59]** 你喜欢历史。可以回顾一下早期 TPU 时代吗？我们也喜欢历史。嗯，

**[00:12:06]** 我想回到2013年，因为那时你们开始自研芯片，比现在的AI军备竞赛早了近十年。

**[00:12:14]** 比Transformer论文早四年。Google为什么要这么做？这个起因是什么？

**[00:12:22]** 事实上，2013年非常神奇，如果你回到当时，所有聪明的工程师都会说自研芯片是个坏主意。CPU，

**[00:12:28]** 总是越来越快，只要等就行。但我们发现，Google面临最大挑战之一是翻译，

**[00:12:34]** 语言翻译至今都很难。我们现在靠 TPUs 和各种模型有很大进步，但要实现像

![关键帧 00:12:34](./assets/20260424_PJQPMv8TqLA_754s.jpg)

> 视觉分析：两人对话场景，背景显示 TPU 8t 图案。

**[00:12:46]** “从一种语言翻译到另一种”这样的应用，第一步就是语音识别。不是只

**[00:12:56]** 做文本，还要理解用户说什么，像星际迷航里的通用翻译器。语音识别

**[00:13:06]** 算力消耗极大，非常重复，矩阵乘法很多。

**[00:13:14]** 我们发现，如果用CPU做这些矩阵乘法——通用计算——

**[00:13:20]** 需要再建两个甚至三个Google公司，才能让每位Google用户每天用语音交互30秒。2013年我们已很

**[00:13:25]** 大，基础设施丰富，但仅为一个场景就三倍扩容，无法承受。Jeff Dean指出，

**[00:13:30]** 我们需要再建一个Google，甚至两个，才能满足需求。但我们发现，如果用定制芯片，可以

**[00:13:37]** 高效提升100倍。数学上完全站得住，CPU要提升100倍，

**[00:13:42]** 要等很久。一般说，等几年CPU会升级，但100倍提升真要等很久，尤其当时CPU性能提升趋势开始放缓，

![关键帧 00:13:42](./assets/20260424_PJQPMv8TqLA_822s.jpg)

> 视觉分析：幻灯片标题：“Our eighth generation TPUs”，背景为 TPU 图片，三位演讲者互动。

**[00:13:53]** 所以很合理。2013年这决定很有争议，TPU v1可以说是传统ASIC，TPU v8早已不是ASIC，那么你会如何

**[00:14:02]** 总结这一转变？

**[00:14:09]** 没错，TPUv1 是ASIC。Google曾不是硬件公司，现在是否算硬件公司也难说，但2013年绝不是硬件公司。所以这是个逐步推进的过程，逐步提升自身能力，每年

**[00:14:14]** 发现需要新功能，需要液冷，需要自定义网络互连等。所以没有

**[00:14:20]** 明确的转折点从ASIC转变。我们不断挑战自我，现在一年推出两款芯片，功能越来越丰富，这也是市场需求。

![关键帧 00:14:20](./assets/20260424_PJQPMv8TqLA_860s.jpg)

> 视觉分析：无实际意义内容。

**[00:14:32]** 只有这样才能让产品成功。你们是怎么开启这段旅程？

**[00:14:39]** 明明发现需要再造一个Google，大多数科技公司都是采购芯片，

**[00:14:47]** 很少有人说我们要自造芯片、自建数据中心、自建机架、自发电。这种做法在Google文化根深蒂固，

**[00:14:57]** 我加盟时深感共鸣。我也要给 Larry 和 Sergey 点赞。这种对“不可能”的健康敬畏很重要。

**[00:15:07]** 想想 Google Maps，谁会想到要自己开飞机、开车去全球测绘？乍一看肯定不合理，

**[00:15:13]** 结果还真合理。所以愿意去尝试“不可能”可能有意义。我们有很多

**[00:15:21]** 传统敢于押注可能失败的事物，现在我们庆祝TPUs的成功，

**[00:15:28]** 也要看到曾经失败的多项尝试。我们不会点名那些没成功的项目。我们做过十小时的历史节目，这很正常。所以，

**[00:15:35]** 这���化好的一点是，有些人有洞见，

**[00:15:41]** 有激情、有勇气去挑战。我们支持他们做，失败了也不会受处罚或影响职业发展。你总能学到东西。大部分时候我们因为失败

**[00:15:50]** 反而让现有产品或别的产品变得更好，

**[00:16:01]** 因为从失败中学到很多。所以很大程度上，敢于下注的文化。最初的投入很小，

**[00:16:13]** 也许是 30 到 40 人负责 TPUv1。也许你不要引用我，但十几百万美元，投入在当时很

**[00:16:24]** 大且有争议。但失败成本低，成功成本却很高。从理解角度来看，

**[00:16:30]** Google 在AI领域最有意思的一点是它集前沿模型、扩展基础设施、

**[00:16:42]** 云服务和顶尖AI芯片公司于一身。感觉这些事情同在一家真的有

**[00:16:47]** 优势。Google这种组合只有它有。能否举些具体例子说明团队如何协作，

**[00:16:54]** 发挥这种整合优势？最重要的是，没有深度团队协作，这些TPU根本做不出来，

**[00:17:00]** 比如语言翻译、语音识别等服务，

**[00:17:05]** 首款TPU其实就是专为一款应用打造的应用级芯片，

**[00:17:16]** 如果专为某应用做硬件，必须深度协作。

**[00:17:24]** 中间几代，以及今天这些TPUs，也被用于广告推荐系统，

**[00:17:32]** 深度协作：你需要什么？瓶颈在哪里？

**[00:17:42]** 如何让硬件更适合你的服务场景？我们要做什么软件栈支持你？最近，

**[00:17:56]** 可以说，没有 DeepMind 深度协作就没有这些TPUs。要准确预测两三年后的未来，

**[00:18:02]** 需要与 DeepMind 团队及研究团队肩并肩合作，他们会实际投入使用。TPUv8有什么例子是你们向DeepMind提案但被

**[00:18:12]** 指出能否这样做？比如他们说：“能否这样调整？”一个特别值得骄傲的点，

**[00:18:18]** 我们没怎么谈但很快会公布，先在这里提前透露，就是TPU8实现了

**[00:18:24]** 显著的延迟下降。我们意识到芯片互连的方式——网络拓扑——这讲起来技术性有点强，

**[00:18:30]** 我尽量让大家容易理解。我们的默认互连方式，不支持低延迟，它只支持

**[00:18:35]** 吞吐量、带宽。它非常擅长大量数据传输。

**[00:18:41]** 但在智能体时代，真正关键的是延迟，即数据通过的最短时间。我们协作开发了一个创新网络拓扑，

**[00:18:48]** 连接芯片，大幅度减少网络直径。芯片之间距离大幅下降。

**[00:18:53]** 这种洞见只有整合团队才能得到，如果只做芯片，就是“让上代更快”。

**[00:18:58]** 当然让上代更快是好事，但根据服务需求调整架构才是真正巨变。两年前

**[00:19:04]** 开发设计时甚至没人谈AI智能体，无法为此做专用芯片设计。

**[00:19:14]** 现在很难想象，两年前还没人谈智能体。好在我们内部早就在讨论智能体和

**[00:19:25]** 实时强化学习等需求，虽然当时不是主流负载。也有其他需求竞争。

**[00:19:31]** 但我们能预见，未来智能体和低延迟会成为核心。我们需要下注，

**[00:19:43]** 因为我们有数千人包括研究员预判未来。如果他们中

**[00:19:51]** 有不少顶级人才，尤其 DeepMind 的团队——我有偏见但他们是全球顶级头脑之一，还有很多其他团队——

**[00:19:59]** 如果他们都说未来两三年会走到这一步，你就必须认真倾听。你刚才说了些相关内容，但我们可以

**[00:20:09]** 多聊聊两款芯片。一方面是双倍工作量，

**[00:20:15]** 这是战略大决策。当时你们做决定时，计算需求还是以训练为主。你们如何考虑这种投入？

**[00:20:21]** 我喜欢研究历史，最大原因之一是研究历史能帮助预测未来。其实我们当时已经有很多 Google 历史作为后盾。针对网页搜索，

**[00:20:30]** 2000年最大的难题是建网页索引，也就是训练模型。那会我们

**[00:20:41]** 并不叫模型训练，只是建索引。是大规模、一次性的工作。你们还记得吗，有个时期

**[00:20:50]** 每隔几个月就会宣布新的网页搜索索引上线。

**[00:21:00]** Google每几个月就会更新底部列出的索引页数量。

**[00:21:12]** 新索引有多少参数，当然其实是页面数。没错，

**[00:21:18]** 新的网页索引上线了。我们几个月勤奋工作，投入几万台服务器（当时算巨量），几万台服务器爬网建索引。索引的价值在哪？

**[00:21:25]** 在于服务它。到2005年甚至2010年，肯定是2005年，我们主要靠基础设施服务索引。

**[00:21:31]** 看到25年前的事，如今都在更短时间内发生。我们知道训练是主流负载，但真正价值在于

**[00:21:39]** 服役模型。自然，训练模型也很难，需要付出巨大努力，但服务���型才是

**[00:21:49]** 为 Gemini Enterprise、搜索、广告、YouTube等创造价值的地方。由此我们押注，看到推理需求呈指数级爆发，预计未来两三年——即使

**[00:21:54]** 现在是2024年——今年我们就要推两款芯片。有趣的是，我想知道你如何思考TPUs的应用场景和受众，

**[00:22:02]** 因为我们都在Google Cloud Next现场，这场发布其实也完全可以在I/O现场。毕竟这

**[00:22:09]** 是Google各部门的基础。你如何看待TPU在Google

**[00:22:16]** 全业务中的角色？我们，团队很幸运，既属于 Cloud 又属于Google整体。我们承担

**[00:22:26]** 双重角色，也把这种担责当成一种特权。我们驱动 Google，所有消费类服务也驱动 Cloud，

![关键帧 00:22:26](./assets/20260424_PJQPMv8TqLA_1346s.jpg)

> 视觉分析：无实际意义内容。

**[00:22:33]** 各类企业服务。用例方面，可以说过去两年Google搜索的变化比过去25年都根本性。

**[00:22:39]** 看看AI模式、AI概要，几乎每次搜索都有

**[00:22:48]** AI模式（或者AI概要）。我每次搜索都在想

**[00:22:54]** 哇这个背后成本多高。我们很自豪做了大量效率优化，

**[00:22:59]** 使得几乎每一次Google搜索都能支撑AI，甚至很多

**[00:23:09]** 不直接创收的搜索也是如此。这种

**[00:23:15]** 设定好的一点是，搜索作为业务单元，它目标并不是赚钱，而是追求输出最高质量、最有吸引力的结果。我很喜欢这一点，

![关键帧 00:23:15](./assets/20260424_PJQPMv8TqLA_1395s.jpg)

> 视觉分析：无实际意义内容。

**[00:23:22]** 有个独立组织也很棒、聚集了最聪明的人负责广告，

**[00:23:27]** on and everything that was happening 25 years ago 
is happening but it's happening on much compressed   time scales. So we knew yes training that's the 
dominant workload but the value is going to come

**[00:23:37]** from serving the model. Of course, all the hard 
work is going to go into building the model in in   a lot of ways, but serving it is where the value 
is created for Gemini Enterprise and search and

**[00:23:46]** ads and YouTube and so much else, right? And so we 
could make a bet that given and we were starting

**[00:23:52]** to see the exponential takeoff for serving and 
if you carry it forward two, three years, okay,

**[00:23:57]** now 2024, this is the year where we're going to go 
after two chips. So, it's funny. Um, I'm curious

![关键帧 00:23:57](./assets/20260424_PJQPMv8TqLA_1437s.jpg)

> 视觉分析：无实际意义内容。

**[00:24:06]** how you think about the use cases for TPUs and 
who it's for because we're all sitting here at

**[00:24:14]** Google Cloud Next. This seems like it could just 
as easily be at IO. I mean, this is this is an

**[00:24:20]** underpinning of Google across every group. How 
how do you think about where TPU is across the

**[00:24:28]** Google portfolio? We so uh my team we're fortunate 
in that we actually uh sit in cloud and we sit in

**[00:24:35]** Google as a whole. So we have a dual role and we 
take this um role and sort of responsibility uh as

**[00:24:40]** a privilege honestly. So we power Google, we power 
all of our consumer services, we power uh cloud

**[00:24:46]** and of course all our our enterprise services. So 
use cases again I I think that it's fair to say

**[00:24:52]** that Google search has changed more fundamentally 
in the last two years than the previous 25.

**[00:24:59]** Right? If you look at AI mode, AI overviews and 
practically every search I run these days has

**[00:25:04]** AI mode which you know I I think about or sorry AI 
overviews. Yeah. Yeah. And I think about just like

**[00:25:10]** wow what is the cost that is going into into that 
and we're very proud of the efficiency work that

**[00:25:16]** we've done to actually make it possible to be able 
to power that into every essentially every Google

**[00:25:23]** search right and and that even non-monetizable 
ones not exactly for us and this is the great

**[00:25:29]** thing about how we're set up is that search's 
goal as a business unit and I love this about   the company is not to make money is to produce the 
highest quality most engaging results. There's a

**[00:25:40]** separate organization, a fantastic one with the 
smartest people as well, that generates the ads,

**[00:25:46]** but search's job is whether it's monetizable or 
not, not not material, produce the best results

**[00:25:52]** and they're doing so with TPUs. Now, on the flip 
side, I uh you'll hear this example tomorrow. We

**[00:25:57]** covered it actually at the session today, so 
I I think I can u mention it here. Citadel,

**[00:26:03]** leading securities trading company in the world, 
they use TPUs. Are they doing it for frontier

**[00:26:08]** model training or model serving? No. But actually 
using TPUs, they've made their trading systems

**[00:26:15]** two to four times more efficient. They've reduced 
cost 30%. Because so it turns out that while it's

![关键帧 00:26:15](./assets/20260424_PJQPMv8TqLA_1575s.jpg)

> 视觉分析：幻灯片标题：“Our eighth generation TPUs”，三位演讲者继续交流。

**[00:26:22]** a specialized architecture, these TPUs over eight 
generations have become more and more general

**[00:26:28]** purpose taking on a number of mathematically 
intense use cases. Wow. as you solve some of

**[00:26:35]** the bottlenecks in the two variants of this chip 
and you're starting to feel like wow this is like

**[00:26:42]** this is an unbelievably high performance piece of 
silicon where are the remaining bottlenecks where

**[00:26:47]** will the issues start to show up elsewhere in the 
chain so one of the big ones for us is reliability

**[00:26:53]** so when we talk about our systems I think at this 
point it's fair to say it's not just 9,600 chips

**[00:26:59]** that are working on a problem in many cases it's 
tens of thousands and dare I say more that are all

**[00:27:05]** coordinating together at literally nancond scale. 
What this means is that if any one chip fails,

**[00:27:13]** the computation stops. I mean literally there's 
like a nervous system where everyone does a bit

**[00:27:18]** of computation and then they exchange information 
with everyone else. I mean I'm I'm simplifying you

**[00:27:23]** all know this is in training in training but also 
in inference. M the scale is smaller in infants

**[00:27:29]** to be fair but we still run on perhaps hundreds 
of chips. What this means is that if any one of

**[00:27:34]** them stop that nervous system stops distributing 
the information the whole computation comes to

**[00:27:40]** a halt. Now still I I assume this was a solved 
problem. No that's actually primary problem for

**[00:27:45]** us and for essentially the entire industry. So we 
are what we're proud of is these fail and these

**[00:27:52]** chips they're wonders of nature and you all are uh 
very knowledgeable here but the they're using the

**[00:27:58]** most advanced physics lithography materials and 
these are the biggest chips essentially to first

**[00:28:05]** order and you can see one one chip is different 
size than the other but these are essentially   the biggest chips that you can buy. What what 
this means is they fail. Yeah. Okay. And and

**[00:28:15]** ours fails other people's fails they fail. there's 
just inherent failure. So, and they fail more than

**[00:28:21]** other chips. And that what that means is if you 
say, "Okay, let's make they're reliable." But if   there are 100,000 of these who are going after a 
piece of computation, what's the chance that one

**[00:28:29]** of them fails? Well, you do the math and it turns 
out pretty good. Se several times a day, I won't

**[00:28:35]** say whether but within that order, several times 
a day, at least one chip is going to fail. Now,

**[00:28:40]** how quickly do you detect that failure? What's 
the needle in the hail stack? How quickly do you

**[00:28:46]** reconfigure the system to start running again? 
If you if you imagine that it takes let's say

**[00:28:51]** you have a failure every hour. If it takes 
you an hour to find the problem and fix it,

**[00:28:57]** you have zero throughput, right? By the time that 
you start up again, you're failing something else.

**[00:29:03]** And if a human being has to get involved in that 
failure detection and diagnosis, there's a rule

**[00:29:09]** human gets in the loop minimum 30 minutes that 
I mean smartest most knowledgeable person that's

**[00:29:15]** that's just what it's going to take minimum and 
it could easily go to ours. So actually our work   in driving this reliability to massive levels we 
are able to deliver on 10,000 chips more than 97%

![关键帧 00:29:15](./assets/20260424_PJQPMv8TqLA_1755s.jpg)

> 视觉分析：幻灯片标题：“Our eighth generation TPUs”，自左至右的三人对话场景，观众视角。

**[00:29:28]** goodput. What that means is actually how much 
forward progress we're making. Not just okay,   what's the theoretical flops if everything 
were going perfect? You said goodput. Good

**[00:29:36]** putut. Never heard that before. Yeah. So it's not 
just throughput because you can be doing a bunch   of computation and not making any progress. 
You could be spinning your wheels. Good putut

**[00:29:45]** you're making progress. 97% means 97 100% is no 
failure. Everything's working perfect. 97% means

![关键帧 00:29:45](./assets/20260424_PJQPMv8TqLA_1785s.jpg)

> 视觉分析：无实际意义内容。

**[00:29:53]** that you're finding those issues when they happen 
and they happen and you're recovering from them   super quickly. Right? And so reliability key key 
issue and I'll just throw one more in. The worst

**[00:30:04]** ones are not fail stop but when you have what 
we call silent data corruption when you have

**[00:30:12]** one chip that silently gets the computation wrong 
every once in a while. Those are the worst. Those

**[00:30:18]** are the because like you know it's like the 
the genius who can get it right almost all the   time. Well if you make one error an hour actually 
that's a big problem because these chips are all

**[00:30:28]** talking to one another. one chips error goes to 
everybody. So it the the problem of making these

**[00:30:35]** things run in production at scale massively 
challenging. It becomes much less a uh narrow

**[00:30:42]** chip development problem and this like very broad 
systems engineering. Exactly. Exactly. And to end

**[00:30:48]** and there's literally hundreds of issues that 
have to be discovered and fixed. each of them

**[00:30:54]** pretty challenging until you get to that point 
where you can deliver this massive reliability   at scale. Well, maybe as as we wrap up here, 
um you know, we're here at Cloud Next, have

**[00:31:06]** this incredible announcement, two chips today. Um 
tomorrow. Tomorrow 5 a.m. 5 a.m. tomorrow. Sorry.

**[00:31:13]** Sorry. Tomorrow. Incredible announcement tomorrow. 
Um what can you tell us about the future looking

**[00:31:19]** a few years out? I'm sure your road maps for 
many generations are already coming together.

**[00:31:24]** Yeah. So the future I think is one of uh really 
re-examining the conventional wisdom. So like   I I think one of the things I'll say though is 
with agentic computing, let me there'll be two

**[00:31:34]** parts to this. CPUs are going to make a comeback. 
That's that's a prediction I'll make here today.

**[00:31:39]** In other words, there's a lot of general purpose 
compute that is involved in running these agents.   They're orchestrating all this inference that 
is taking place. They're creating sandboxes,

**[00:31:48]** virtual machines to uh build code, run it, check 
the results, and then find the next set of uh

**[00:31:54]** outputs. So general purpose compute is going to 
make a comeback. But so that's one side. The other

**[00:32:00]** side is the age of specialization is going to 
continue. So I wouldn't be surprised if two chips

**[00:32:07]** here, we're going to find additional workloads 
that might need their own chip. Whether it's us,

**[00:32:14]** I'm going to make a prediction for industry, not 
for Google. But at at a time when general purpose

**[00:32:19]** CPUs are really only improving in performance 5% a 
year, normalize the cost. You have to specialize,

![关键帧 00:32:19](./assets/20260424_PJQPMv8TqLA_1939s.jpg)

> 视觉分析：无实际意义内容。

**[00:32:27]** right? If you're really going to go after 
these brand new workloads, so two chips might   become more. Maybe they're not TPUs, maybe 
they're going after a different workload,

**[00:32:34]** but the age of specialization will be upon us 
in the future. Well, I mean, that is a great

**[00:32:40]** place to leave this. Thank you so much. Yep. 
Thank you. It's a lot of fun. Appreciate it.
