---
格式版本: 2
标题: "ClusterAttention: A training-free speedup of bidirectional attention"
原文链接: "https://arxiv.org/abs/2608.26965"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 27 Aug 2026 11:04:41 UTC (6,310 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-30T15:09:50+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-30T15:08:58+08:00"
入库时间: "2026-08-30T07:09:50.377Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "latency"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 40
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是无需训练的稀疏注意力算法，通过递归聚类和质心补偿加速TabPFN及视频生成模型，并非超节点、AI Rack或可独立复用的机架级通信、网络、调度、内存或RAS基础设施。当前页面为可追溯的arXiv论文摘要页，披露2—6倍、1.8倍加速及至少99%精度等实验结果，但缺少完整工程正文、生产平台和商业部署证据。固定知识库未见ClusterAttention同项记录，但Top 5不完整，只能确认本文自述的新方法与实验结果，不能据此认定历史首次。命中“应用与模型效率”强否决项，未命中超节点高价值准入通道。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-30T15:10:06+08:00"
AI主题相关性: 1
AI来源权威性: 11
AI新颖性: 14
AI技术细节: 6
AI商业部署信号: 0
AI完整性: 8
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.55"
AI评分知识库SHA256: "ca56afec1ff616b1fa4394cd1d08b1e2a6af47eccbe70c802c9f1fea6037e803"
AI评分知识库检索词: "[\"GPU\",\"NPU\",\"PDF\",\"arxiv.org/pdf/2608.26965\",\"HTML\",\"arxiv.org/html/2608.26965v1\",\"v1\",\"GPUs\",\"TabPFN-3\",\"arxiv.org/abs/2605.13986\",\"T2V\",\"arxiv.org/abs/2503.20314\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"PDF\",\"HTML\",\"v1\"],\"rank\":-16.98578930816261},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\"],\"rank\":-10.52985569710448},{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"PDF\",\"HTML\"],\"rank\":-10.335917217367548},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"PDF\",\"v1\",\"GPUs\"],\"rank\":-9.11507456480313},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"HTML\"],\"rank\":-8.684846200756855}]"
AI摘要: "ClusterAttention 提出一种免训练的双向注意力加速方法，通过快速递归聚类生成块稀疏注意力，并用质心补偿减少误差。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-30T18:09:01.132Z"
采集批次: "2026年8月30日14点11分37秒"
采集批次ID: "20260830-141137-130"
去重键: "https://arxiv.org/abs/2608.26965"
---

## Computer Science > Machine Learning

## Title:ClusterAttention: A training-free speedup of bidirectional attention

Authors:[Kasper Nordenram](https://arxiv.org/search/cs?searchtype=author&query=Nordenram,+K), [Amelie Dittmann](https://arxiv.org/search/cs?searchtype=author&query=Dittmann,+A)

[View PDF](https://arxiv.org/pdf/2608.26965) [HTML (experimental)](https://arxiv.org/html/2608.26965v1)

> Abstract:This paper introduces ClusterAttention, a general training-free speedup of bidirectional attention layers. Existing sparse attention methods either rely on structure in the input, such as order in language or spatial proximity in images, or use slow clustering processes amortized over several forward passes. ClusterAttention instead uses a fast recursive clustering method that adapts to the geometry of the keys and queries in each attention head to produce useful clusters. This method allows setting the size of the clusters arbitrarily. We utilize this by setting all clusters to be a fixed size that is a power of two, allowing the block-sparse attention to run at the same latency per query-key interaction as dense attention on GPUs. We also derive an expression for the output error in sparse attention, that explains the counterintuitive experimental finding that tight clusters can lead to larger errors than random clusters. We then derive the error when excluded clusters are compensated through their centroids, and show that this error shrinks with tighter clusters. We integrate this compensation into the method.  
> On large-scale tabular data ClusterAttention speeds up TabPFN-3 [arXiv:2605.13986](https://arxiv.org/abs/2605.13986) by two to six times, while retaining at least 99% of the dense accuracy. To our knowledge, it is the first training-free method that can be successfully applied in the setting of unstructured input and a single forward pass. For video generation with Wan 2.1-14B T2V [arXiv:2503.20314](https://arxiv.org/abs/2503.20314), ClusterAttention achieves output closer to dense attention and a larger speedup (1.8x versus 1.4x) compared to SVOO [arXiv:2603.18636](https://arxiv.org/abs/2603.18636), a leading method developed specifically for this domain, both run without offline calibration.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | [arXiv:2608.26965](https://arxiv.org/abs/2608.26965) \[cs.LG\] |
|  | (or [arXiv:2608.26965v1](https://arxiv.org/abs/2608.26965v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.26965](https://doi.org/10.48550/arXiv.2608.26965) |
| Related DOI: | [https://doi.org/10.5281/zenodo.22118033](https://doi.org/10.5281/zenodo.22118033) |

## Submission history

From: Kasper Nordenram \[[view email](https://arxiv.org/show-email/5c33ae38/2608.26965)\]  
**\[v1\]** Thu, 27 Aug 2026 11:04:41 UTC (6,310 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.26965) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
