---
格式版本: 2
标题: "Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy"
原文链接: "https://arxiv.org/abs/2609.00817"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 1 Sep 2026 07:17:17 UTC (893 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T18:51:34+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T18:51:29+08:00"
入库时间: "2026-09-02T10:51:34.871Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI%20Rack&searchtype=all"
匹配关键词:
  - "AI Rack"
  - "AI"
  - "Scale-up"
  - "NVL72"
  - "performance"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 54
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是大规模AI训练在scale-up、scale-out及rack-scale配置下的扩展性和多租户干扰研究，来自arXiv原创预印本页面。相较知识库已有GB300 NVL72性能与部署信息，新增点是覆盖最高2400 GPU、六类集群、五种并行策略及噪声模型的跨平台研究设计；但当前正文仅有摘要，未披露具体性能结果、通信开销、拓扑、带宽或端到端指标，也无客户、量产或部署动作。命中“仅摘要壳/事实不足”否决项，当前页面本身不足以作为完整技术信息源，封顶54分。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T18:51:51+08:00"
AI主题相关性: 18
AI来源权威性: 11
AI新颖性: 13
AI技术细节: 8
AI商业部署信号: 0
AI完整性: 4
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"AI Rack\",\"GB300\",\"NVL72\",\"rack-scale\",\"GPU\",\"PDF\",\"arxiv.org/pdf/2609.00817\",\"HTML\",\"arxiv.org/html/2609.00817v1\",\"HPC\",\"v1\",\"GPUs\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0089\",\"title\":\"Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GB300\",\"NVL72\",\"rack-scale\",\"GPU\",\"HTML\",\"GPUs\"],\"rank\":-15.41055293804854},{\"id\":\"runtime-dccd09f98e2ca028070793f0\",\"title\":\"OCI Achieves NVIDIA Exemplar Cloud Validation for NVIDIA GB300 NVL72 and HGX B300 | cloud-infrastructure\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-22\",\"matchedTerms\":[\"GB300\",\"NVL72\",\"rack-scale\",\"GPU\",\"GPUs\"],\"rank\":-14.88223690050492},{\"id\":\"july-correct-0062\",\"title\":\"Why AI Racks Need an Open Signal Conditioning Standard\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AI Rack\",\"NVL72\",\"rack-scale\",\"GPU\",\"PDF\",\"v1\"],\"rank\":-13.83583103492094},{\"id\":\"runtime-0fa5f6513e1d75f79b9b89e6\",\"title\":\"https://www.datacenterknowledge.com/ai-data-centers/ai-rack-density-s-real-limits-power-cooling-failure-risk\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-28\",\"matchedTerms\":[\"AI Rack\",\"GB300\",\"NVL72\",\"GPU\",\"PDF\",\"HTML\",\"GPUs\"],\"rank\":-13.778182409842689},{\"id\":\"july-correct-0012\",\"title\":\"Microsoft to deploy AMD Helios rackscale solution to support inference workloads and Azure services\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AI Rack\",\"rack-scale\",\"GPU\",\"HPC\",\"GPUs\"],\"rank\":-12.927522215864622}]"
AI摘要: "该研究系统评估了多租户环境下大规模AI训练的可扩展性与性能，覆盖最高2400块GPU，对比五种并行策略在不同超算集群与互连配置下的通信开销。文中还构建了噪声模型，分析并发训练任务间的干扰，为多租户场景的性能行为提供了关键洞察。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T17:13:35.652Z"
采集批次: "2026年9月2日18点28分48秒"
采集批次ID: "20260902-182848-894"
去重键: "https://arxiv.org/abs/2609.00817"
---

## Computer Science > Distributed, Parallel, and Cluster Computing

## Title:Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy

[View PDF](https://arxiv.org/pdf/2609.00817) [HTML (experimental)](https://arxiv.org/html/2609.00817v1)

> Abstract:Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads and their impact across different interconnects by evaluating scale-up, scale-out, and rack-scale configurations under multiple allocation schemes. Finally, we study how multiple concurrent training jobs interfere with each other by designing a realistic noise model. We design a benchmark suite of AI models to evaluate the performance of five distinct parallelisation strategies across different supercomputing clusters, including Alps, Leonardo, LUMI, JUPITER, NVL72 GB300, and DGX A100. Our work provides a systematic characterization of the scalability and execution efficiency of distributed AI training, while offering key insights into performance behavior under realistic multi-tenant scenarios.

| Comments: |  |
| --- | --- |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC) |
| Cite as: | [arXiv:2609.00817](https://arxiv.org/abs/2609.00817) \[cs.DC\] |
|  | (or [arXiv:2609.00817v1](https://arxiv.org/abs/2609.00817v1) \[cs.DC\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.00817](https://doi.org/10.48550/arXiv.2609.00817) |

## Submission history

From: Jacopo Raffi \[[view email](https://arxiv.org/show-email/f742bf74/2609.00817)\]  
**\[v1\]** Tue, 1 Sep 2026 07:17:17 UTC (893 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.00817) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
