---
格式版本: 2
标题: "FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving"
原文链接: "https://arxiv.org/abs/2608.19758"
发布日期: "2026-08-20"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 20 Aug 2026 08:02:55 UTC (617 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-21T16:15:01+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-21T16:14:13+08:00"
入库时间: "2026-08-21T08:15:01.510Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=NVIDIA&searchtype=all"
匹配关键词:
  - "deployment"
  - "performance"
相关厂家:
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 20
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为LLM长上下文预填充优化，属于算法层，仅测试平台提及NVIDIA H20 GPU，与超节点/AI Rack/机柜级AI基础设施及其供电散热互连等主题无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-21T16:15:10+08:00"
AI主题相关性: 0
AI来源权威性: 8
AI新颖性: 5
AI技术细节: 2
AI商业部署信号: 0
AI完整性: 5
AI摘要: "FlashPrefill V2 提出一种面向长上下文 LLM 服务的块稀疏预填充注意力机制，在 FlashPrefill 基础上引入均值校正项抑制近似误差，并重设计算子以对齐 FlashAttention-3/4。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:12:21.577Z"
采集批次: "2026年8月21日13点55分53秒"
采集批次ID: "20260821-135553-743"
去重键: "https://arxiv.org/abs/2608.19758"
---

## Computer Science > Computation and Language

## Title:FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Authors:[Qihang Fan](https://arxiv.org/search/cs?searchtype=author&query=Fan,+Q), [Huaibo Huang](https://arxiv.org/search/cs?searchtype=author&query=Huang,+H), [Zhiying Wu](https://arxiv.org/search/cs?searchtype=author&query=Wu,+Z), [Bingning Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+B), [Ran He](https://arxiv.org/search/cs?searchtype=author&query=He,+R)

[View PDF](https://arxiv.org/pdf/2608.19758) [HTML (experimental)](https://arxiv.org/html/2608.19758v1)

> Abstract:Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2608.19758](https://arxiv.org/abs/2608.19758) \[cs.CL\] |
|  | (or [arXiv:2608.19758v1](https://arxiv.org/abs/2608.19758v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.19758](https://doi.org/10.48550/arXiv.2608.19758) |

## Submission history

From: Qihang Fan \[[view email](https://arxiv.org/show-email/e5899eb6/2608.19758)\]  
**\[v1\]** Thu, 20 Aug 2026 08:02:55 UTC (617 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.19758) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
