---
格式版本: 2
标题: "Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs"
原文链接: "https://arxiv.org/abs/2609.00575"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:local:strict_original_body"
发布时间证据: "**[\\[v1\\]]( Tue, 1 Sep 2026 02:11:40 UTC (1,153 KB)"
发布时间校准原因: "arXiv论文正文明确标注v1提交时间为2026年9月1日，该时间即为文章首次发布时间。"
发布时间校准置信度: "1"
发布时间候选数量: 23
发布时间严格候选数量: 7
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:35:56+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 6744
发现时间: "2026-09-04T01:32:14+08:00"
入库时间: "2026-09-03T17:36:03.295Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 18
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为MoE大模型压缩的学术论文，仅提及GPU内存需求，未涉及超节点、AI Rack、机柜级系统、供电散热互连或量产落地等核心主题，与项目关注范围无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:37:50+08:00"
AI主题相关性: 3
AI来源权威性: 8
AI新颖性: 4
AI技术细节: 2
AI商业部署信号: 0
AI完整性: 1
AI摘要: "PARSER 是一种用于压缩混合专家大语言模型的新残差稀疏化方法，将压缩目标从最小化单个矩阵误差转为保留专家输出误差。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:17:56.368Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.00575"
---

## Computer Science > Artificial Intelligence

## Title:Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

[View PDF](https://arxiv.org/pdf/2609.00575) [HTML (experimental)](https://arxiv.org/html/2609.00575v2)

> Abstract:Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative compression technique that decomposes each projection matrix of an expert into a shared base matrix and per-expert residual matrix, and then compresses the residuals. Existing sparsification methods compress each residual matrix independently by minimizing its compression error, thereby minimizing the error of each projection matrix. However, our analysis shows that this objective is misaligned with preserving model accuracy after compression. In an expert, the final output is produced through computations coupled across multiple projections and hidden representations. Therefore, even small errors in individual matrices can propagate through hidden representations and projection interactions, leading to large expert output errors and accuracy degradation. To address this misalignment, we propose PARSER, a new residual sparsification method that shifts the compression objective from minimizing isolated matrix errors to preserving the expert output error. PARSER achieves this by introducing output importance, which measures the actual contribution to the expert output error. Our experiments show that, compared with existing methods, PARSER narrows the accuracy gap to the uncompressed model by 1.41 $\times$ on Qwen and 1.44 $\times$ on DeepSeek, while achieving the same peak memory reduction. Our code is available at [this https URL](https://github.com/OSSS-KU/PARSER).

| Comments: |  |
| --- | --- |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2609.00575](https://arxiv.org/abs/2609.00575) \[cs.AI\] |
|  | (or [arXiv:2609.00575v2](https://arxiv.org/abs/2609.00575v2) \[cs.AI\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00575](https://doi.org/10.48550/arXiv.2609.00575) |

## Submission history

From: Seungwoo Jung \[[view email](https://arxiv.org/show-email/76054ca8/2609.00575)\]  
**[\[v1\]](https://arxiv.org/abs/2609.00575v1)** Tue, 1 Sep 2026 02:11:40 UTC (1,153 KB)  
**\[v2\]** Wed, 2 Sep 2026 04:21:16 UTC (1,153 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00575) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
