---
格式版本: 2
标题: "Squeezing More from Limited Data with Recursive Transformers"
原文链接: "https://arxiv.org/abs/2608.26973"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 11:18:27 UTC (1,229 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:13:11+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:11:38+08:00"
入库时间: "2026-08-30T07:13:11.164Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 31
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是有限训练数据下的递归Transformer与因子化嵌入，新增事实为三个递归模型在10M和100M词预算下优于标准Transformer，但仅呈现论文摘要。arXiv预印本具有一定学术可追溯性，尚非正式同行评审成果；固定知识库未显示同一研究，但不能据此认定首次出现。文章不涉及超节点、AI机柜、机架级互连、供电、液冷或部署，也无客户、量产和商业信号，命中“应用与模型效率/单一模型训练”硬否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:13:19+08:00"
AI主题相关性: 0
AI来源权威性: 11
AI新颖性: 10
AI技术细节: 4
AI商业部署信号: 0
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"Scale-up\",\"RAS\",\"C3\",\"BClbahar\",\"PDF\",\"arxiv.org/pdf/2608.26973\",\"HTML\",\"arxiv.org/html/2608.26973v1\",\"v1\",\"arxiv.org/abs/2608.26973\",\"arxiv.org/abs/2608.26973v1\",\"doi.org/10.48550/arXiv.2608.26973\"]"
AI评分知识库命中: "[{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-12.808968674720035},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\"],\"rank\":-11.333006229301258},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"RAS\",\"C3\",\"PDF\",\"v1\"],\"rank\":-8.011504167715488},{\"id\":\"runtime-5581f72bc66fda8e6c6bf711\",\"title\":\"Advancing Standards-Based AI Fabric RAS Through COSMOS Integration\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-10\",\"matchedTerms\":[\"RAS\",\"PDF\"],\"rank\":-7.089762522059228},{\"id\":\"july-correct-0115\",\"title\":\"锚定 300kW 整机柜演进方向 OAII 社区三项规范联合发布，树立AIDC基础设施建设标准答案\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"v1\"],\"rank\":-6.992347863111062}]"
AI摘要: "Serdar Gülbahar等人提出用递归Transformer改善有限数据下的预训练：在10M和100M词预算下，共享块递归结构和因子化嵌入让模型超越标准Transformer。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:08:54.527Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.26973"
---

## Computer Science > Computation and Language

## Title:Squeezing More from Limited Data with Recursive Transformers

Authors:[Serdar Gülbahar](https://arxiv.org/search/cs?searchtype=author&query=G%C3%BClbahar,+S), [Lukas Edman](https://arxiv.org/search/cs?searchtype=author&query=Edman,+L), [Alexander Fraser](https://arxiv.org/search/cs?searchtype=author&query=Fraser,+A)

[View PDF](https://arxiv.org/pdf/2608.26973) [HTML (experimental)](https://arxiv.org/html/2608.26973v1)

> Abstract:Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens. We study this behavior across 10M-100M word pre-training budgets, two corpora, and multiple downstream evaluations, and find that optimal size depends strongly on both the data budget and the downstream target. We argue that standard Transformers scale down poorly to this setting, because embeddings consume a large fraction of the parameter budget and per-token computation is tied to representational capacity. To address this coupling, we study recursive Transformers, reusing a shared block across depth to scale compute, together with factorized embeddings to reduce vocabulary-map parameters. We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners.

| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| --- | --- |
| Cite as: | [arXiv:2608.26973](https://arxiv.org/abs/2608.26973) \[cs.CL\] |
|  | (or [arXiv:2608.26973v1](https://arxiv.org/abs/2608.26973v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.26973](https://doi.org/10.48550/arXiv.2608.26973) |

## Submission history

From: Serdar Gülbahar \[[view email](https://arxiv.org/show-email/8f10e447/2608.26973)\]  
**\[v1\]** Thu, 27 Aug 2026 11:18:27 UTC (1,229 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.26973) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
