---
格式版本: 2
标题: "Learning how to Forget: Fine-tuning for Long-Context Sparse Attention"
原文链接: "https://arxiv.org/abs/2608.19920"
发布日期: "2026-08-20"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 20 Aug 2026 11:37:04 UTC (141 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-21T16:15:06+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-21T16:14:13+08:00"
入库时间: "2026-08-21T08:15:06.876Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=NVIDIA&searchtype=all"
匹配关键词:
  - "GPU"
  - "AI"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 15
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为长上下文稀疏注意力微调，仅提及单卡A100作硬件示例，未涉及超节点、AI Rack、机柜级基础设施、互连、供电或液冷等任何项目主题，与关注范围无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-21T16:15:18+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 5
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 5
AI摘要: "该文提出一种面向稀疏注意力KV缓存策略的模型微调方法，适用于任意缓存策略，仅在单块NVIDIA A100 40GB GPU上即可运行，使模型与策略共同适应，常优于采用精确注意力的基线模型。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:12:19.012Z"
采集批次: "2026年8月21日13点55分53秒"
采集批次ID: "20260821-135553-743"
去重键: "https://arxiv.org/abs/2608.19920"
---

## Computer Science > Computation and Language

**arXiv:2608.19920** (cs)

## Title:Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

[View PDF](https://arxiv.org/pdf/2608.19920) [HTML (experimental)](https://arxiv.org/html/2608.19920v1)

> Abstract:A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues ([this https URL](https://github.com/awslabs/keys_values)), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.

| Comments: | 39 pages, no figures |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2608.19920](https://arxiv.org/abs/2608.19920) \[cs.CL\] |
|  | (or [arXiv:2608.19920v1](https://arxiv.org/abs/2608.19920v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.19920](https://doi.org/10.48550/arXiv.2608.19920) Focus to learn more  arXiv-issued DOI via DataCite (pending registration) |

## Submission history

From: Matthias Seeger \[[view email](https://arxiv.org/show-email/96e5ff84/2608.19920)\]  
**\[v1\]** Thu, 20 Aug 2026 11:37:04 UTC (141 KB)

## Bibliographic and Citation Tools

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

## Code, Data and Media Associated with this Article

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com/))*

DagsHub *([What is DagsHub?](https://dagshub.com/))*

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

## Demos

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai/))*

## Recommenders and Search Tools

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

## arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.19920) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
