---
格式版本: 2
标题: "Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs"
原文链接: "https://arxiv.org/abs/2609.04168"
发布日期: "2026-09-03"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Thu, 3 Sep 2026 17:53:44 UTC (482 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-06T21:23:37+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-06T21:21:19+08:00"
入库时间: "2026-09-06T13:23:37.959Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "performance"
  - "latency"
  - "throughput"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 12
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文主题为SoC上ML计算图的算子并行优化，属于边缘端嵌入式系统，与超节点/AI Rack/机柜级AI基础设施、高速互连、供电液冷等完全无关。仅因SoC含GPU命中关键词，无有效商业或技术细节，判定无关。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-06T21:23:43+08:00"
AI主题相关性: 2
AI来源权威性: 8
AI新颖性: 2
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "Para-Pipe提出一种面向异构SoC的分层映射框架，在流水线架构中集成算子级并行，以平衡推理吞吐与延迟并降低处理器间通信开销。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-06T23:44:26.879Z"
采集批次: "2026年9月6日19点27分09秒"
采集批次ID: "20260906-192709-237"
去重键: "https://arxiv.org/abs/2609.04168"
---

## Computer Science > Distributed, Parallel, and Cluster Computing

## Title:Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

[View PDF](https://arxiv.org/pdf/2609.04168) [HTML (experimental)](https://arxiv.org/html/2609.04168v1)

> Abstract:As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM [this http URL](http://big.little/) CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution.

| Comments: |  |
| --- | --- |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF) |
| Cite as: | [arXiv:2609.04168](https://arxiv.org/abs/2609.04168) \[cs.DC\] |
|  | (or [arXiv:2609.04168v1](https://arxiv.org/abs/2609.04168v1) \[cs.DC\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.04168](https://doi.org/10.48550/arXiv.2609.04168) |
| Journal reference: | IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 44, no. 12, pp. 4472-4485, Dec. 2025 |
| Related DOI: | [https://doi.org/10.1109/TCAD.2025.3568348](https://doi.org/10.1109/TCAD.2025.3568348) |

## Submission history

From: Yujie Zhang \[[view email](https://arxiv.org/show-email/4718c30c/2609.04168)\]  
**\[v1\]** Thu, 3 Sep 2026 17:53:44 UTC (482 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.04168) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
