---
格式版本: 2
标题: "APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization"
原文链接: "https://arxiv.org/abs/2608.25380"
发布日期: "2026-08-27"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:local:strict_original_body"
发布时间证据: "Published Time: Thu, 27 Aug 2026 00:25:05 GMT"
发布时间校准原因: "正文明确标注发布时间为8/27，优先于提交时间8/26。"
发布时间校准置信度: "1"
发布时间候选数量: 12
发布时间严格候选数量: 3
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-27T23:25:34+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 5969
发现时间: "2026-08-27T23:21:22+08:00"
入库时间: "2026-08-27T15:25:55.899Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "latency"
  - "AI"
相关厂家:
  - "NVIDIA"
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
图片摘要:
  - "✗ ./assets/img-5838d115.png | decorative | 与正文无关的基金赞助商logo"
  - "✗ ./assets/img-26707af3.png | decorative | 与正文无关的基金赞助商logo"
  - "✗ ./assets/img-43a00bd2.png | decorative | 与正文无关的基金赞助商logo"
AI优质: "否"
AI打分: 40
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是面向高分辨率扩散Transformer的专用加速器，通过注意力概率引导剪枝、量化及软硬件协同提升单一模型工作负载效率，并非超节点或机架级AI基础设施。arXiv论文已注明被ICCAD 2026接收，来源可追溯；固定知识库未见APT相关历史事实，但Top 5不完整，不能据此认定首次出现。本文新增APDT、TAFA、双精度计算单元及相对A100最高8.16×加速、14.98×能效等论文实验结果，但没有生产平台、客户、量产或部署证据，且当前页面主要提供摘要而非完整论文正文。命中“应用与模型效率”及“学术原型缺少生产采用”否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-27T23:26:09+08:00"
AI主题相关性: 2
AI来源权威性: 12
AI新颖性: 13
AI技术细节: 6
AI商业部署信号: 0
AI完整性: 7
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.12"
AI评分知识库SHA256: "13aade27529fda213430b6def2f11efe60b599f8676ae127cad35798cf7169c8"
AI评分知识库检索词: "[\"Scale-up\",\"NVIDIA\",\"Google\",\"APT\",\"URL\",\"arxiv.org/abs/2608.25380\",\"GMT\",\"HTML\",\"PDF\",\"arxiv.org/pdf/2608.25380\",\"arxiv.org/html/2608.25380v1\",\"SOTA\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0104\",\"title\":\"NVIDIA Vera Rubin 提升每瓦性能，为全球合作伙伴实现最低 Token 成本\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"NVIDIA\",\"Google\"],\"rank\":-9.639726356741043},{\"id\":\"july-correct-0024\",\"title\":\"Salience Labs Wants To Scale Up AI With Silicon Photonics Optical Switch\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"NVIDIA\",\"Google\",\"URL\"],\"rank\":-8.603913890182213},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"URL\",\"HTML\",\"PDF\"],\"rank\":-8.359283396509767},{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"NVIDIA\",\"Google\"],\"rank\":-8.243499350218729},{\"id\":\"july-correct-0033\",\"title\":\"Microsoft, Alphabet, Meta Pivot from Buy to Build in AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NVIDIA\",\"Google\",\"URL\"],\"rank\":-7.521737504176961}]"
AI摘要: "该论文提出APT，一种软硬件协同设计的扩散Transformer加速器，以注意力概率为统一重要性指标，联合优化剪枝与量化。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T18:59:26.972Z"
采集批次: "2026年8月27日21点49分30秒"
采集批次ID: "20260827-214930-609"
去重键: "https://arxiv.org/abs/2608.25380"
---

Title: APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

URL Source: https://arxiv.org/abs/2608.25380

Published Time: Thu, 27 Aug 2026 00:25:05 GMT

Markdown Content:
[Skip to main content](https://arxiv.org/abs/2608.25380#content)[](https://arxiv.org/IgnoreMe)[Search](https://arxiv.org/search)[Submit](https://arxiv.org/user/create)[Donate](https://info.arxiv.org/about/donate.html)[Log in](https://arxiv.org/login)

Search arXiv 

 Press Enter to search · [Advanced search](https://arxiv.org/search/advanced)

# Computer Science > Hardware Architecture

**arXiv:2608.25380** (cs) 

 [Submitted on 26 Aug 2026]

# Title:APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

Authors:[Sungyeob Yoo](https://arxiv.org/search/cs?searchtype=author&query=Yoo,+S), [Seeyeon Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+S), [Joonyong Park](https://arxiv.org/search/cs?searchtype=author&query=Park,+J), [Seunghee Han](https://arxiv.org/search/cs?searchtype=author&query=Han,+S), [Joo-Young Kim](https://arxiv.org/search/cs?searchtype=author&query=Kim,+J)

View a PDF of the paper titled APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization, by Sungyeob Yoo and 4 other authors

[View PDF](https://arxiv.org/pdf/2608.25380)[HTML (experimental)](https://arxiv.org/html/2608.25380v1)
> Abstract:Recent advances in generative AI have significantly increased the demand for high-resolution image and video generation, positioning diffusion models as a core technology. Among them, Diffusion Transformers (DiTs) have emerged as the state-of-the-art (SOTA) models due to their scalability and output quality. However, self-attention in DiTs incurs significant computational overhead, leading to excessively long latency as the complexity grows with the fourth power of the output resolution. While prior works have attempted to mitigate this cost using sparsity and quantization techniques, they fall short of effectively reducing the computational cost in high-resolution DiTs. 
> 
> In this paper, we present APT, a software-hardware co-designed accelerator for high-resolution DiTs. APT leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling. At the algorithm level, we propose Attention Probability-guided Adaptive Dual Thresholding (APDT), which dynamically performs element selection and precision assignment using dual thresholds. To ensure compatibility with memory-efficient FlashAttention, we introduce Timestep-Aware FlashAttention (TAFA), which predicts attention probabilities across timesteps by exploiting temporal similarity. At the architecture level, we co-design a specialized accelerator that efficiently supports irregular sparsity and dual-precision execution, featuring dynamic mask management, address translation, dual-precision compute units, and a tile-based dataflow. Finally, we evaluate APT on SOTA DiT models, including PixArt-α, Stable Diffusion 3, and FLUX. APT achieves up to 8.16×speedup and 14.98×higher energy efficiency over NVIDIA A100, and up to 3.01×speedup and 2.04×higher energy efficiency over EXION, a SOTA diffusion model accelerator.

Comments:9 pages, 17 figures, 3 tables. Accepted to the IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)
Subjects:Hardware Architecture (cs.AR)
Cite as:[arXiv:2608.25380](https://arxiv.org/abs/2608.25380) [cs.AR]
(or [arXiv:2608.25380v1](https://arxiv.org/abs/2608.25380v1) [cs.AR] for this version)
[https://doi.org/10.48550/arXiv.2608.25380](https://doi.org/10.48550/arXiv.2608.25380)

Focus to learn more

 arXiv-issued DOI via DataCite (pending registration)
Related DOI:[https://doi.org/10.1145/3831252.3834102](https://doi.org/10.1145/3831252.3834102)

Focus to learn more

 DOI(s) linking to related resources

## Submission history

 From: Sungyeob Yoo [[view email](https://arxiv.org/show-email/9737ea6c/2608.25380)] 

**[v1]** Wed, 26 Aug 2026 05:08:40 UTC (2,212 KB)

[](https://arxiv.org/abs/2608.25380)Full-text links:
## Access Paper:

 View a PDF of the paper titled APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization, by Sungyeob Yoo and 4 other authors

*   [View PDF](https://arxiv.org/pdf/2608.25380)
*   [HTML (experimental)](https://arxiv.org/html/2608.25380v1)
*   [TeX Source](https://arxiv.org/src/2608.25380)

[view license](http://creativecommons.org/licenses/by/4.0/ "Rights to this article")

### Current browse context:

cs.AR

[<prev](https://arxiv.org/prevnext?id=2608.25380&function=prev&context=cs.AR "previous in cs.AR (accesskey p)") | [next>](https://arxiv.org/prevnext?id=2608.25380&function=next&context=cs.AR "next in cs.AR (accesskey n)")

[new](https://arxiv.org/list/cs.AR/new) | [recent](https://arxiv.org/list/cs.AR/recent) | [2026-08](https://arxiv.org/list/cs.AR/2026-08)

 Change to browse by: 

[cs](https://arxiv.org/abs/2608.25380?context=cs)

### References & Citations

*   [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2608.25380)
*   [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2608.25380)
*   [Semantic Scholar](https://api.semanticscholar.org/arXiv:2608.25380)

export BibTeX citation Loading...

## BibTeX formatted citation

×

Data provided by: [](https://arxiv.org/abs/2608.25380)

### Bookmark

Bibliographic Tools 

# Bibliographic and Citation Tools

- [x] Bibliographic Explorer Toggle 

Bibliographic Explorer _([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))_

- [x] Connected Papers Toggle 

Connected Papers _([What is Connected Papers?](https://www.connectedpapers.com/about))_

- [x] Litmaps Toggle 

Litmaps _([What is Litmaps?](https://www.litmaps.co/))_

- [x] scite.ai Toggle 

scite Smart Citations _([What are Smart Citations?](https://www.scite.ai/))_

Code, Data, Media 

# Code, Data and Media Associated with this Article

- [x] alphaXiv Toggle 

alphaXiv _([What is alphaXiv?](https://alphaxiv.org/))_

- [x] Links to Code Toggle 

CatalyzeX Code Finder for Papers _([What is CatalyzeX?](https://www.catalyzex.com/))_

- [x] DagsHub Toggle 

DagsHub _([What is DagsHub?](https://dagshub.com/))_

- [x] GotitPub Toggle 

Gotit.pub _([What is GotitPub?](http://gotit.pub/faq))_

- [x] Huggingface Toggle 

Hugging Face _([What is Huggingface?](https://huggingface.co/huggingface))_

- [x] ScienceCast Toggle 

ScienceCast _([What is ScienceCast?](https://sciencecast.org/welcome))_

Demos 

# Demos

- [x] Replicate Toggle 

Replicate _([What is Replicate?](https://replicate.com/docs/arxiv/about))_

- [x] Spaces Toggle 

Hugging Face Spaces _([What is Spaces?](https://huggingface.co/docs/hub/spaces))_

- [x] Spaces Toggle 

TXYZ.AI _([What is TXYZ.AI?](https://txyz.ai/))_

Related Papers 

# Recommenders and Search Tools

- [x] Link to Influence Flower 

Influence Flower _([What are Influence Flowers?](https://influencemap.cmlab.dev/))_

- [x] Core recommender toggle 

CORE Recommender _([What is CORE?](https://core.ac.uk/services/recommender))_

*   [Author](https://arxiv.org/abs/2608.25380)
*   [Venue](https://arxiv.org/abs/2608.25380)
*   [Institution](https://arxiv.org/abs/2608.25380)
*   [Topic](https://arxiv.org/abs/2608.25380)

 About arXivLabs  

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.25380) | [Disable MathJax](javascript:setMathjaxCookie()) ([What is MathJax?](https://info.arxiv.org/help/mathjax.html)) 

 We gratefully acknowledge support from our **major funders**, [**member institutions**](https://info.arxiv.org/about/ourmembers.html), , and all contributors. 

[About](https://info.arxiv.org/about)·[Help](https://info.arxiv.org/help)·[Contact](https://info.arxiv.org/help/contact.html)·[Subscribe](https://info.arxiv.org/help/subscribe)·[Copyright](https://info.arxiv.org/help/license/index.html)·[Privacy](https://info.arxiv.org/help/policies/privacy_policy.html)·[Accessibility](https://info.arxiv.org/help/web_accessibility.html)·[Operational Status (opens in new tab)](https://status.arxiv.org/)

Major funding support from
