---
格式版本: 2
标题: "SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers"
原文链接: "https://arxiv.org/abs/2609.01343"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 14:52:57 UTC (1,743 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:36+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:39+08:00"
入库时间: "2026-09-02T10:52:36.564Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "performance"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 39
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是SMELT MoE循环Transformer的模型架构与计算缩放规律研究，并非超节点或机架级AI基础设施。arXiv预印本来源可追溯但未经正式同行评审；固定知识库未发现该方案的既有记录，本文新增了最高54B参数实验、计算预算匹配及训练FLOPs节省6.8%—18.0%等结果，但不能据此确认历史首次。技术信息集中于模型层循环、损失与注意力机制，不涉及机架拓扑、互连、供电、散热或RAS，也无客户、量产和部署信号。命中“应用与模型效率”硬否决项，且当前页面主要为摘要，不能进入业务情报库。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:54:04+08:00"
AI主题相关性: 2
AI来源权威性: 10
AI新颖性: 14
AI技术细节: 6
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"Scale-up\",\"SMELT\",\"PDF\",\"arxiv.org/pdf/2609.01343\",\"HTML\",\"arxiv.org/html/2609.01343v1\",\"FLOPs\",\"KV\",\"arxiv.org/abs/2609.01343\",\"arxiv.org/abs/2609.01343v1\",\"doi.org/10.48550/arXiv.2609.01343\",\"arxiv.org/show-email/f17ebe4f/2609.01343\"]"
AI评分知识库命中: "[{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\"],\"rank\":-10.41584664390761},{\"id\":\"runtime-44415c412091149e8ca9b29f\",\"title\":\"三年磨一剑，重新发明HBM - 智东西\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"Scale-up\",\"KV\"],\"rank\":-6.976576428821429},{\"id\":\"july-correct-0117\",\"title\":\"ODCC分享 | 华为严可荣：OCS全光超节点集群技术研究分析\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"KV\"],\"rank\":-6.828764547460016},{\"id\":\"runtime-4abbccc42d96af674efc7768\",\"title\":\"OpenAI’ Jalapeño: Better Than Nvidia Blackwell\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-25\",\"matchedTerms\":[\"Scale-up\",\"FLOPs\",\"KV\"],\"rank\":-6.758965034157779},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\"],\"rank\":-6.721112210816699}]"
AI摘要: "研究者在混合专家变压器上提出SMELT配方，让中间层循环两次，并在匹配FLOPs、参数和KV缓存预算下缩放至54B参数，发现其损失下降更快，可在最优计算前沿节省6.8%–18.0%训练FLOPs，该优势在下游任务（尤其代码任务）中依然成立。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:23.902Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.01343"
---

## Computer Science > Machine Learning

## Title:SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Authors:[Shaowen Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+S), [Ge Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+G), [Kairong Luo](https://arxiv.org/search/cs?searchtype=author&query=Luo,+K), [Yuhao Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Y), [Shaofan Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+S), [Jiaheng Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+J), [Wenhao Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+W), [Shen Yan](https://arxiv.org/search/cs?searchtype=author&query=Yan,+S), [Jian Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+J)

[View PDF](https://arxiv.org/pdf/2609.01343) [HTML (experimental)](https://arxiv.org/html/2609.01343v1)

> Abstract:Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | [arXiv:2609.01343](https://arxiv.org/abs/2609.01343) \[cs.LG\] |
|  | (or [arXiv:2609.01343v1](https://arxiv.org/abs/2609.01343v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.01343](https://doi.org/10.48550/arXiv.2609.01343) |

## Submission history

From: Shaowen Wang \[[view email](https://arxiv.org/show-email/f17ebe4f/2609.01343)\]  
**\[v1\]** Tue, 1 Sep 2026 14:52:57 UTC (1,743 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.01343) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
