---
格式版本: 2
标题: "Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis"
原文链接: "https://arxiv.org/abs/2609.00746"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 05:24:37 UTC (322 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:19+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:52:19.319Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 20
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是视觉语言模型微调导致文本能力退化及 attention-sink 诊断，新增 Sink Strength 指标、六组模型对比和若干负面实验结果；arXiv 原始预印本可追溯，但当前页面仅提供摘要。固定知识库未见该论文重复记录，但其新增事实属于模型适配研究，不涉及超节点、AI Rack、机架级互连、供电、散热、RAS或部署。无客户、量产、交付及机架级商业信号，命中应用与模型效率类强否决，当前页面不值得进入超节点业务情报库。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:53:10+08:00"
AI主题相关性: 0
AI来源权威性: 11
AI新颖性: 4
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 5
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"PDF\",\"arxiv.org/pdf/2609.00746\",\"HTML\",\"arxiv.org/html/2609.00746v1\",\"v1\",\"LLM\",\"VLM\",\"VL\",\"VLM-LLM\",\"QK-RMSNorm\",\"arxiv.org/abs/2609.00746\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\",\"LLM\"],\"rank\":-13.73935445447137},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\",\"v1\",\"VL\"],\"rank\":-10.630171179861968},{\"id\":\"july-correct-0115\",\"title\":\"锚定 300kW 整机柜演进方向 OAII 社区三项规范联合发布，树立AIDC基础设施建设标准答案\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"v1\",\"LLM\"],\"rank\":-9.525234386891393},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"PDF\",\"v1\",\"LLM\"],\"rank\":-9.365959513879563},{\"id\":\"july-correct-0131\",\"title\":\"DeepSeek-V4如何在昇腾超节点高效完成全参数后训练？SLAI T-Rex技术报告解读\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"LLM\"],\"rank\":-6.919263704682261}]"
AI摘要: "研究发现，将预训练大语言模型微调为视觉-语言模型会损伤其文本能力，尤其在严格遵循输出规则的任务上，原因在于视觉微调破坏了注意力汇聚位置。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:27.622Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00746"
---

## Computer Science > Machine Learning

## Title:Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

Authors:[Minsik Choi](https://arxiv.org/search/cs?searchtype=author&query=Choi,+M), [Geewook Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+G), [Young Geun Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+Y+G)

[View PDF](https://arxiv.org/pdf/2609.00746) [HTML (experimental)](https://arxiv.org/html/2609.00746v1)

> Abstract:Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that require following exact output rules, such as instruction following, chain-of-thought reasoning graded on a strictly parsed final answer, and similar evaluations with strict graders. We trace this gap to attention-sink corruption: VL fine-tuning perturbs the early sink position that anchors a large fraction of attention probability, and how well the base LLM preserves its sink tracks how much of the affected capability survives adaptation. Building on this view, we introduce Sink Strength, a single scalar computed on the base LLM in a few seconds on a single GPU that predicts post-VL degradation without any VL training. It consistently tracks relative degradation across the six VLM-LLM pairs and multiple format-sensitive tasks. Complementing this diagnostic, we find that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training. These negative results underscore the value of screening backbones with Sink Strength before VL training and narrow the intervention space toward head-selective training-time protection.

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | [arXiv:2609.00746](https://arxiv.org/abs/2609.00746) \[cs.LG\] |
|  | (or [arXiv:2609.00746v1](https://arxiv.org/abs/2609.00746v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00746](https://doi.org/10.48550/arXiv.2609.00746) |

## Submission history

From: Minsik Choi \[[view email](https://arxiv.org/show-email/7a44b41c/2609.00746)\]  
**\[v1\]** Tue, 1 Sep 2026 05:24:37 UTC (322 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00746) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
