---
格式版本: 2
标题: "LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism"
原文链接: "https://arxiv.org/abs/2609.00857"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 07:55:45 UTC (9,394 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:52:08+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:36+08:00"
入库时间: "2026-09-02T10:52:09.062Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=GPU&searchtype=all"
匹配关键词:
  - "GPU"
  - "Scale-up"
  - "performance"
  - "latency"
  - "bandwidth"
  - "throughput"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 45
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是面向LLM推理的片上IMC-NoC软硬件协同架构，而非机架级AI基础设施。arXiv原始预印本提出LEAP，将IMC PE、NMC PE和INC分工，并采用预填充/解码解耦与动态PE重组，报告吞吐提升至少1.52倍、能效提升24.91倍；固定知识库未显示同一方案重复，但未命中不能证明首次出现。当前页面仅有摘要，未提供真实生产平台、端到端部署、客户、量产或主流采用证据。命中学术原型缺少生产验证及应用/模型推理导向的强否决项，且摘要壳完整性有限。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:52:50+08:00"
AI主题相关性: 6
AI来源权威性: 10
AI新颖性: 13
AI技术细节: 12
AI商业部署信号: 0
AI完整性: 4
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"GPU\",\"LLM\",\"IMC-NoC\",\"PDF\",\"arxiv.org/pdf/2609.00857\",\"HTML\",\"IMC\",\"PE\",\"arxiv.org/html/2609.00857v1\",\"LEAP\",\"NMC\",\"INC\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"LLM\",\"PDF\",\"HTML\",\"PE\"],\"rank\":-8.916694782784841},{\"id\":\"runtime-7aadc05d024a3a525d014ae9\",\"title\":\"MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-24\",\"matchedTerms\":[\"GPU\",\"PDF\",\"PE\",\"NMC\",\"INC\"],\"rank\":-8.730933513301695},{\"id\":\"runtime-569639751a0dbe7ef3ffccc5\",\"title\":\"[2608.17503] Predict Before Replay: Joint FEC and Flight Control for Reliable Scale-Up Links\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-18\",\"matchedTerms\":[\"PDF\",\"HTML\",\"PE\"],\"rank\":-8.032611962777274},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"PDF\",\"HTML\",\"PE\",\"INC\"],\"rank\":-7.671653124801715},{\"id\":\"july-correct-0015\",\"title\":\"AMD to join the optical interconnect party with 2027 Instinct GPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"LLM\",\"PE\",\"LEAP\",\"INC\"],\"rank\":-7.401718054862884}]"
AI摘要: "该文提出一个名为LEAP的软硬件协同架构，统一集成IMC、NMC与INC，分别处理静态权重、动态数据和部分结果归约，并配合分区、映射与调度框架及prefill-decode分离部署，用于提升LLM推理性能。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:30.429Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00857"
---

## Computer Science > Hardware Architecture

## Title:LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

Authors:[Yimin Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y), [Yue Jiet Chong](https://arxiv.org/search/cs?searchtype=author&query=Chong,+Y+J), [Xuanyao Fong](https://arxiv.org/search/cs?searchtype=author&query=Fong,+X)

[View PDF](https://arxiv.org/pdf/2609.00857) [HTML (experimental)](https://arxiv.org/html/2609.00857v1)

> Abstract:LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data generated during run-time. Furthermore, the massive number of parameters in LLM necessitates scale-up architectures where on-chip data movement is often the primary performance bottleneck. This article presents a hardware-software co-design framework that unifies distributed compute, memory, and communication into a seamless processing-communication fabric. On the hardware side, we propose a scalable architecture, named LEAP, that integrates IMC PE, NMC PE, and INC. This allows each hardware layer to execute specialized tasks: IMC for static weights, NMC for dynamic data, and INC for partial result reduction. On the software side, we introduce a partitioning, mapping, and scheduling framework optimized for key metrics in LLM serving, including throughput and latency. To address the distinct computational intensities of the prefill and decode phases, we present a prefill-decode disaggregation approach that dynamically reconfigures PE organizations to maximize resource utilization. Compared to commercial GPU platforms, the proposed architecture provides a throughput and an energy efficiency improvement of $\geq{}1.52\times$ and $24.91\times$, respectively.

| Comments: |  |
| --- | --- |
| Subjects: | Hardware Architecture (cs.AR) |
| Cite as: | [arXiv:2609.00857](https://arxiv.org/abs/2609.00857) \[cs.AR\] |
|  | (or [arXiv:2609.00857v1](https://arxiv.org/abs/2609.00857v1) \[cs.AR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00857](https://doi.org/10.48550/arXiv.2609.00857) |

## Submission history

From: Yimin Wang \[[view email](https://arxiv.org/show-email/c13c83fc/2609.00857)\]  
**\[v1\]** Tue, 1 Sep 2026 07:55:45 UTC (9,394 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00857) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
