---
格式版本: 2
标题: "StrataCL: Fabric-Native Communication Library for Production Supernodes"
原文链接: "https://arxiv.org/abs/2607.26444"
发布日期: "2026-07-29"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 29 Jul 2026 03:45:05 UTC (641 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-07-30T13:44:10+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-07-30T13:44:01+08:00"
入库时间: "2026-07-30T05:44:10.345Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=supernode&searchtype=all"
匹配关键词:
  - "supernode"
  - "NPU"
  - "bandwidth"
  - "throughput"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Direct URL Open"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "是"
AI打分: 92
AI分档: "高置信优质"
AI质检状态: "通过"
AI打分理由: "论文直接针对生产超节点提出通信库，在华为CloudMatrix384实测，含架构、性能数据，新型且高相关性。"
AI质检模型: "deepseek-v4-flash"
AI质检时间: "2026-07-30T13:44:50+08:00"
AI主题相关性: 20
AI来源权威性: 12
AI新颖性: 20
AI技术细节: 20
AI商业部署信号: 10
AI完整性: 10
采集批次: "2026年7月30日10点43分13秒"
采集批次ID: "20260730-104313-925"
去重键: "https://arxiv.org/abs/2607.26444"
---

## Computer Science > Distributed, Parallel, and Cluster Computing

## Title:StrataCL: Fabric-Native Communication Library for Production Supernodes

Authors:[Tiancheng Hu](https://arxiv.org/search/cs?searchtype=author&query=Hu,+T), [Jin Qin](https://arxiv.org/search/cs?searchtype=author&query=Qin,+J), [Yuzheng Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Y), [Ke Liu](https://arxiv.org/search/cs?searchtype=author&query=Liu,+K), [TangShengsheng Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+T), [Sheng Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+S), [Zhongzhe Hu](https://arxiv.org/search/cs?searchtype=author&query=Hu,+Z), [Tianlun Hu](https://arxiv.org/search/cs?searchtype=author&query=Hu,+T), [Wei Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+W), [Lijun Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+L), [Jingbin Zhou](https://arxiv.org/search/cs?searchtype=author&query=Zhou,+J), [Xiaoming Bao](https://arxiv.org/search/cs?searchtype=author&query=Bao,+X), [Hongwei Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+H), [Jieru Zhao](https://arxiv.org/search/cs?searchtype=author&query=Zhao,+J), [Huimin Cui](https://arxiv.org/search/cs?searchtype=author&query=Cui,+H), [Tao Xie](https://arxiv.org/search/cs?searchtype=author&query=Xie,+T), [Chenxi Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+C)

[View PDF](https://arxiv.org/pdf/2607.26444) [HTML (experimental)](https://arxiv.org/html/2607.26444v1)

> Abstract:Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because user and communication buffers are managed separately, causing redundant data copies or costly user-buffer registration. This paper presents StrataCL, a zero-redundancy and fabric-native communication library for production supernodes. StrataCL introduces registration-on-allocation to realize user-buffer direct communication, and designs communication operators with workload-balanced NPU-core partitioning and NPU-driven SDMA offloading to exploit supernode architecture features. On the Huawei CloudMatrix384, StrataCL improves collective bus bandwidth by up to 1.6x and improves MoE dispatch/combine bus bandwidth by up to 1.4x. Across three production workloads, StrataCL improves LLM inference throughput by 1.9x, reduces P99 TTFT by 2.2x, and reduces LLM and Recsys training iteration time by 1.4x and 1.3x, respectively.

| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC) |
| --- | --- |
| Cite as: | [arXiv:2607.26444](https://arxiv.org/abs/2607.26444) \[cs.DC\] |
|  | (or [arXiv:2607.26444v1](https://arxiv.org/abs/2607.26444v1) \[cs.DC\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2607.26444](https://doi.org/10.48550/arXiv.2607.26444) |

## Submission history

From: Tiancheng Hu \[[view email](https://arxiv.org/show-email/2397c53c/2607.26444)\]  
**\[v1\]** Wed, 29 Jul 2026 03:45:05 UTC (641 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2607.26444) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
