--- 格式版本: 2 标题: "Deep|Optics: Scale-In Moves Optics Inside the Tray, and the Bandwidth Opportunity Could 10x Scale-Up" 原文链接: "https://fundaai.substack.com/p/deepoptics-scale-in-moves-optics" 发布日期: "2026-09-01" 发布时间校准状态: "found" 发布时间需复核: "否" 发布时间来源: "llm:scrape:strict_html_body" 发布时间证据: "div class=pencraft pc-display-flex pc-gap-12 pc-alignItems-center pc-reset byline-wrapper: Sep 01, 2026" 发布时间校准原因: "正文附近byline-wrapper明确显示发布时间为Sep 01, 2026,属标题附近标注的发布时间,优先级高于元数据。" 发布时间校准置信度: "1" 发布时间候选数量: 15 发布时间严格候选数量: 3 发布时间原页读取状态: "source template page reused from URL open" 发布时间未找到原因: "" 发布时间校准时间: "2026-09-02T01:48:26+08:00" 发布时间仲裁状态: "confirmed" 发布时间仲裁尝试次数: 1 发布时间仲裁耗时毫秒: 6213 发现时间: "2026-09-02T01:46:39+08:00" 入库时间: "2026-09-01T17:49:16.834Z" 来源平台: "Substack 数据中心相关博客搜索" 搜索渠道: "source_template" 搜索词: "site:substack.com Compute tray" 匹配关键词: - "GPU" - "Scale-up" - "bandwidth" - "Compute tray" - "XPU" - "HBM" - "NVLink" - "UALink" - "CXL" - "Vera Rubin" - "Marvell" - "SRAM" - "deployment" - "AI" 相关厂家: - "NVIDIA" - "Meta" - "Microsoft" - "Google" 相关专家: [] 内容类型: "网页" 抓取工具: "Free Fetch + Defuddle" 清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取" 原始附件: [] AI优质: "是" AI打分: 81 AI分档: "高置信优质" AI质检状态: "通过" AI打分理由: "深度分析光互连进入tray内scale-in层,覆盖die-to-die/die-to-memory带宽、路线图及NVIDIA/Marvell/Google投资与部署信号,技术细节充实,与AI Rack机柜级互连高度相关,但来源为个人博客…" AI质检模型: "zj-deepseek-v4-flash" AI质检时间: "2026-09-02T01:49:30+08:00" AI主题相关性: 18 AI来源权威性: 7 AI新颖性: 16 AI技术细节: 18 AI商业部署信号: 13 AI完整性: 9 AI摘要: "光互连正从机架级“规模扩展”移入托盘内部,即芯片到芯片与芯片到内存的“规模内”链路,其单GPU带宽约为规模扩展的10倍。" AI摘要模型: "ali-deepseek-v4-flash" AI摘要时间: "2026-09-01T23:07:13.430Z" 采集批次: "2026年9月1日23点56分52秒" 采集批次ID: "20260901-235652-761" 去重键: "https://fundaai.substack.com/p/deepoptics-scale-in-moves-optics" --- #### Core conclusion Optics are moving inside the tray, and the memory tier it targets is the largest bandwidth tier in the system, roughly ten times the scale-up bandwidth per GPU and about one hundred times the scale-out. The capital is already moving, with $2 billion each from NVIDIA into Lumentum and Coherent, $3.25 billion from Marvell for Celestial AI, and a revenue ceiling of up to $120 billion on the Google and Marvell warrant. Scale-in, in this note, means die-to-die and chip-to-memory optical links at millimeters to about a meter, which Lumentum’s CEO on August 27 called “the thing to watch over the next 12 to 18 months”. The word is spreading with split meanings since NVIDIA uses it for something else, and Section 3 sorts out who means what. The kicker is Section 7, where our checks point to the largest scale-in exploration running at Google, built around memory pooling, and the warrant’s disclosed scope, “memory interface controllers and near-memory compute”, is what we believe to be its public paper trail. The work is early, and first-generation products start well below the ceiling. #### 1\. Scale-In’s Coming-Out Week At an investment conference on August 27, Lumentum CEO Michael Hurlston said, “I think the thing that is going to be significant is the scale-in, and that’s the thing to watch over the next 12 to 18 months... Are you going to see it actually go in-tray and serve this high bandwidth connectivity between memory and between GPUs?” He was describing optical links inside trays. Scale-in was also featured in imec’s slides on the first day of Semicon Taiwan, in a talk on 3D integrated optics for chip-to-chip interconnect. Also on August 27, Marvell said on its FY27Q2 earnings call that its fiscal 2028 revenue outlook for scale-up optics “has increased meaningfully compared to prior expectations”, against the roughly $300 million category framing it gave a quarter earlier, and that it is on pace to make about $1 billion of capacity prepayments to suppliers in fiscal 2027. The parts are the same since scale-up optics uses the lasers, optical engines, packaging, and test flows that scale-in will need, just at longer reach. Scale-up optics ships starting in 2027 while scale-in follows two to three years later, if not more, because those same parts still have to clear tighter thermal, reliability, test, and other requirements within the tray. On August 19, a week before the earnings call, Google and Marvell announced a partnership that, in the filing’s words, “spans a comprehensive range of custom silicon programs that attach to the TPU ecosystem”. Marvell issued Google a warrant on up to 59 million shares. Most of it vests in 240 equal installments, one for each $500 million of qualifying revenue through fiscal 2033, which is where the $120 billion figure in the press comes from. None of that revenue is committed. The disclosed scope covers AI inference accelerators, storage controllers, network interface controllers, memory interface controllers, and near-memory compute. We believe the last two items refer to the pooled-memory tier behind scale-in, and Section 7 explains why. At Hot Chips 2026, NVIDIA used Scale-In for a DPU services network, and Meta used it for chiplet-level data movement inside its MTIA 400 accelerator. Scale-in now has at least three different technical meanings. #### 2\. The Definition Adopted in This Note: Optical Links Inside the Tray, From Millimeters to About One Meter In this note, scale-in refers to optical links inside the tray that serve high-bandwidth connectivity between memory and between GPUs. This includes die-to-die connections between chiplets and compute die, and chip-to-memory connections such as XPU to HBM, over reaches from millimeters to about one meter. Links from the XPU to pooled memory belong to the adjacent pooled-memory fabric, which this note treats separately. The term also appears as scaling-in and, at Marvell, as scale-inside. AI systems use different interconnect media at different physical layers. Within the package, CoWoS, interposers, and RDL dominate. On the board, systems rely on PCB traces and retimers. Within the rack, copper DAC and AEC cables handle scale-up traffic. Within the data center, pluggable optics dominate, and between data centers, coherent optics dominate. As bandwidth demand per XPU rises, optics enters one layer at a time, and the electrical-to-optical boundary moves progressively closer to the chip. The interconnect ladder from in-package to between-data-centers. Each step inward raises the bandwidth convention: scale-in runs 2-20 TB/s per XPU in-package and 1-8 TB/s in-tray, against 7-29 Tb/s for scale-up in-rack and 0.8-3.2 Tb/s for scale-out. This also explains why bandwidth requirements rise at each inward layer. Scale-out is measured by network-port bandwidth, scale-up by fabric bandwidth, and scale-in approaches memory bandwidth. The closer the link is to the chip, the higher the aggregate bandwidth per XPU and the more demanding the power and density requirements. Comparing those tiers requires conversion: a byte is eight bits, so 1 TB/s equals 8 Tb/s, and a link can be quoted one way or both ways. Lumentum’s ECOC 2025 material puts per-GPU scale-up bandwidth at 7.2, 14.4, and 28.8 Tbps for Blackwell, Rubin, and Feynman, counted one way, while NVIDIA quotes NVLink 6 at 3.6 TB/s, counted both ways. The two agree once directions are matched, since Lumentum’s 14.4 Tbps is per direction while NVIDIA’s 3.6 TB/s is the both-ways total of the same link. This note quotes bandwidth one way throughout, and memory-bus figures are aggregates. The optical boundary moves inward from scale-out to scale-in. The optical boundary moves inward from scale-out to scale-in. Source: FUNDA. #### 3\. Everyone Means the Tray, Except NVIDIA The optics and memory camp uses scale-in for the tier inside the tray. Marvell’s solutions page positions its portfolio “across scale-in, scale-up, scale-out and scale-across architectures”, and its blogs call the inner tier scale-inside, meaning die-to-die interconnects that move data between compute and memory dies inside XPUs, extending to chip-to-chip links within a tray. Avicena defines it as “AI scale-in (die-to-die and die-to-memory connectivity)” with reach from a few millimeters to one meter. Keysight bounds it at “the silicon and package domain including pre-silicon design, chiplets, memory and die-to-die/package interconnects”. Eliyan frames its market as “AI scale-up and scale-in networks”. LightXcelerate, a Palo Alto startup, states that its optical chiplet supports scale-in, scale-up, and scale-out across UCIe, UALink, PCIe, CXL, and NVLink. Executives use the word the same way. Coherent CTO Julie Eng, on a Celesta Capital TechSurge panel with Eliyan founder Ramin Farjadrad and Lumentum cloud CTO Matt Sysak, put scale-across and scale-out in deployment now, large-scale scale-up deployment next year, and scale-in after that, describing it as solving “inter-die interconnects with light” when copper runs out in the die-to-die space. She also noted renewed interest in NRZ signaling for short links, “especially scale-up and die-to-die scale-in”. Eliyan founder Ramin Farjadrad, Coherent CTO Julie Eng, and Lumentum cloud CTO Matt Sysak on the deployment sequence for optics. Eliyan founder Ramin Farjadrad, Coherent CTO Julie Eng, and Lumentum cloud CTO Matt Sysak on the deployment sequence for optics. Source: Celesta Capital TechSurge, Aug 2026. The research world has picked the word up as well. On the opening day of Semicon Taiwan, imec presented a slide at its ITF Taiwan forum titled 3D integrated optics for chip-to-chip scale-in interconnect, showing wide-and-slow area I/O with low-speed SerDes and a step from pluggable optics to co-packaged optics to an optics layer stacked directly under the XPU and HBM. On imec’s chart, the 3D step pulls away from co-packaged optics by roughly two orders of magnitude in bandwidth density per unit of energy (Gbps/mm per pJ/bit) by 2035. That came just four days after the Lumentum and Marvell statements. imec at ITF Taiwan, Semicon Taiwan: 3D integrated optics for chip-to-chip scale-in interconnect. imec at ITF Taiwan, Semicon Taiwan: 3D integrated optics for chip-to-chip scale-in interconnect. Source: imec. Photo by Jukan (@jukan05 on X), August 31, 2026. At Hot Chips 2026, NVIDIA used the same word for a different layer. Its five-network AI factory model pairs Scale-In with the BlueField-4 DPU, a per-node infrastructure network for agentic CPU capacity, storage, security, and orchestration, quantified at 800 Gb/s per Vera Rubin compute tray, plus 4x1.6 Tb/s of scale-out. NVIDIA also has the pooled-memory tier this note discusses, called Context Memory (CMX), built on BlueField-4, and does not call it scale-in. Meta’s MTIA 400 slide deck uses “Scale-in DMA & streaming reductions” for data movement between chiplets inside the package, over electrical links. Microsoft’s Azure documentation has long used scale-in for removing virtual machine instances in autoscaling. The table below explains each term. Meta MTIA 400: scale-in DMA and streaming reductions inside a multi-chiplet design. Meta MTIA 400: scale-in DMA and streaming reductions inside a multi-chiplet design. Source: Meta, Hot Chips 2026. Lumentum, Marvell, Coherent, Avicena, Keysight, Eliyan, LightXcelerate and imec all use scale-in for the tier inside the tray: die-to-die and die-to-memory optical links from a few millimeters to one meter. NVIDIA uses Scale-In for the BlueField-4 DPU services network at 800 Gb/s per Vera Rubin tray, Meta for chiplet-level DMA inside the MTIA 400 package, and Microsoft Azure for removing virtual machine instances in autoscaling. Scale-in as defined in this note, other uses of the term and the adjacent pooled tier. Scale-in as defined in this note, other uses of the term and the adjacent pooled tier. Source: FUNDA, redrawn from public sources. MPU means memory processing unit, covered in Section 7. NVIDIA sells the pooled tier this note describes, under the CMX name. It has invested $2 billion each in Lumentum and Coherent with purchase commitments, and licensed Groq’s technology for SRAM-first inference. Its current scale-up racks run on copper. The five networking infrastructures of the NVIDIA AI factory. Scale-In paired with BlueField-4. The five networking infrastructures of the NVIDIA AI factory. Scale-In paired with BlueField-4. Source: NVIDIA, Hot Chips 2026.