---
格式版本: 2
标题: "What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers"
原文链接: "https://arxiv.org/abs/2608.18399"
发布日期: "2026-08-19"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 19 Aug 2026 00:10:38 UTC (1,368 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-20T14:22:21+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-20T14:19:36+08:00"
入库时间: "2026-08-20T06:22:21.805Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 5
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "实际内容为计算机视觉领域关于Vision Transformer注意力迁移的学术论文，与超节点、AI Rack、机柜级AI基础设施等主题完全无关，仅因搜索词命中Scale-up，不构成有效资料。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-20T14:22:56+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "该研究在ImageNet-100上用ViT-S学生模型测量注意力迁移，发现蒸馏后的注意力结构几乎完美且永久接近教师，但鲁棒性差距随训练成熟可闭合，强制改变注意力结构也无明显响应，结论是分布偏移下的性能缺陷来自特征而非可见注意力结构。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:19:54.356Z"
采集批次: "2026年8月20日14点19分32秒"
采集批次ID: "20260820-141932-079"
去重键: "https://arxiv.org/abs/2608.18399"
---

## Computer Science > Computer Vision and Pattern Recognition

## Title:What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers

Authors:[Jesse Ponnock](https://arxiv.org/search/cs?searchtype=author&query=Ponnock,+J)

[View PDF](https://arxiv.org/pdf/2608.18399) [HTML (experimental)](https://arxiv.org/html/2608.18399v1)

> Abstract:Vision transformers (ViTs) trained to copy a pretrained teacher's attention maps recover most of fine-tuning's in-distribution accuracy yet fall measurably short of it under distribution shift, as recent work has shown. What the copy delivers has never been measured directly in the attention structure and tied to robustness. We build that instrumentation for ViT-S students of a self-supervised teacher on ImageNet-100, and report three findings that triangulate one conclusion. First, the transfer is essentially perfect and permanently so: the distilled student's attention ends up roughly two orders of magnitude closer to the teacher's than fine-tuning does, and does not drift with additional training. Second, the gap is real at 14 $\times$ fewer parameters and 10 $\times$ less data than previously studied, but it has a time axis. It tracks training maturity, and completing the schedules that the stopping rule interrupted closes it below our pre-registered threshold in two of three seeds, with comparisons at equal accuracy giving the same result. The endpoint gap at this scale is substantially a training-maturity artifact: robustness matures later than accuracy, and stopping rules tuned to accuracy undersample it. Third, forcing cross-row redundancy down by half the structural separation between the distilled and fine-tuned conditions produces no detectable robustness response under two registered ways of matching accuracy. Verified transfer, a gap that closes while the structure never moves, and a null under direct intervention are together consistent with the deficit residing in features, not in the visible attention structure. This is elimination plus intervention, and its scope is the regime we measured. In this regime, attention overlays show where a model looks, not what it knows.

| Comments: |  |
| --- | --- |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | [arXiv:2608.18399](https://arxiv.org/abs/2608.18399) \[cs.CV\] |
|  | (or [arXiv:2608.18399v1](https://arxiv.org/abs/2608.18399v1) \[cs.CV\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.18399](https://doi.org/10.48550/arXiv.2608.18399) |

## Submission history

From: Jesse Ponnock \[[view email](https://arxiv.org/show-email/e3fcd6f8/2608.18399)\]  
**\[v1\]** Wed, 19 Aug 2026 00:10:38 UTC (1,368 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.18399) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
