---
格式版本: 2
标题: "PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition"
原文链接: "https://arxiv.org/abs/2609.01024"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 10:21:10 UTC (3,707 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:54+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:39+08:00"
入库时间: "2026-09-02T10:52:55.017Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 40
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是MoE模型推理优化，提出PCoMoE路径组合、分层剪枝和执行引擎，并报告最高1.31倍端到端加速及10%准确率提升；来源为可追溯的arXiv预印本，但当前页面仅有摘要，缺少实验配置与完整工程细节。固定知识库未显示同一方案，但其并非完整历史，本文新增仅能确认该学术框架及论文内实验结果。没有机架级拓扑、互连、RAS、供电或液冷机制，也无生产平台、客户、量产或部署证据，命中“应用与模型效率”及学术原型否决项。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:54:31+08:00"
AI主题相关性: 2
AI来源权威性: 10
AI新颖性: 14
AI技术细节: 8
AI商业部署信号: 0
AI完整性: 6
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"Scale-up\",\"PCoMoE\",\"PDF\",\"arxiv.org/pdf/2609.01024\",\"HTML\",\"arxiv.org/html/2609.01024v1\",\"LLM\",\"v1\",\"URL\",\"github.com/gzyyy0/PCoMoE\",\"arxiv.org/abs/2609.01024\",\"arxiv.org/abs/2609.01024v1\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"HTML\",\"LLM\",\"v1\",\"URL\"],\"rank\":-15.637571845262228},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"HTML\",\"v1\",\"URL\"],\"rank\":-14.83282482837444},{\"id\":\"july-correct-0010\",\"title\":\"Schneider Electric and AMD release first Helios platform reference design to accelerate AI Factory deployment\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"PDF\",\"LLM\",\"v1\",\"URL\"],\"rank\":-10.653736339978341},{\"id\":\"july-correct-0115\",\"title\":\"锚定 300kW 整机柜演进方向 OAII 社区三项规范联合发布，树立AIDC基础设施建设标准答案\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"HTML\",\"LLM\",\"v1\"],\"rank\":-9.525234386891393},{\"id\":\"july-correct-0062\",\"title\":\"Why AI Racks Need an Open Signal Conditioning Standard\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Scale-up\",\"PDF\",\"LLM\",\"v1\",\"URL\"],\"rank\":-9.320559657710064}]"
AI摘要: "PCoMoE提出一种路径组合执行框架，将混合专家（MoE）模型的推理从粗粒度的专家选择转为细粒度的路径组合，并结合分层剪枝与硬件友好执行引擎降低计算冗余。实验显示，该方法可实现最高1.31倍端到端推理加速，同时将模型准确率提升10%。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:20.933Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.01024"
---

## Computer Science > Computation and Language

## Title:PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition

[View PDF](https://arxiv.org/pdf/2609.01024) [HTML (experimental)](https://arxiv.org/html/2609.01024v1)

> Abstract:Mixture-of-Experts (MoE) architectures scale Large Language Model (LLM) capacity efficiently by activating a sparse subset of experts per token. However, modern MoE inference remains heavily constrained by the rigid, whole-expert abstraction. Existing frameworks manage, schedule, or prune experts as atomic execution units, which fixes the optimization boundary too early and leaves fine-grained intra-expert computational redundancy underexplored. In this work, we present PCoMoE, a path-compositional execution framework that shifts MoE inference from coarse-grained expert selection to fine-grained path composition. PCoMoE incorporates a path-level formulation of expert computation, a compatibility-aware layer-wise pruning strategy to suppress low-value path combinations, and a hardware-friendly execution engine to exploit reusable sub-expert structures under strictly bounded overheads. Experimental results demonstrate that PCoMoE achieves up to a 1.31x end-to-end inference speedup while enhancing model accuracy by 10%. The code is available at [this https URL](https://github.com/gzyyy0/PCoMoE)

| Comments: |  |
| --- | --- |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | [arXiv:2609.01024](https://arxiv.org/abs/2609.01024) \[cs.CL\] |
|  | (or [arXiv:2609.01024v1](https://arxiv.org/abs/2609.01024v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.01024](https://doi.org/10.48550/arXiv.2609.01024) |

## Submission history

From: Ziyan Gan \[[view email](https://arxiv.org/show-email/af4b1cf0/2609.01024)\]  
**\[v1\]** Tue, 1 Sep 2026 10:21:10 UTC (3,707 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.01024) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
