---
格式版本: 2
标题: "Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs"
原文链接: "https://arxiv.org/abs/2608.10614"
发布日期: "2026-08-11"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Tue, 11 Aug 2026 07:57:22 UTC (497 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-12T11:24:59+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-12T11:23:15+08:00"
入库时间: "2026-08-12T03:24:59.996Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=throughput&searchtype=all"
匹配关键词:
  - "throughput"
  - "GPU"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 5
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "学术论文内容为MoE模型安全攻击，与超节点/AI Rack/机柜级AI基础设施等硬件架构、供电、散热、互连及量产落地完全无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-12T11:25:13+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
采集批次: "2026年8月12日2点04分20秒"
采集批次ID: "20260812-020420-289"
去重键: "https://arxiv.org/abs/2608.10614"
---

## Computer Science > Cryptography and Security

## Title:Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs

Authors:[Rui Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+R), [Wenbo Jiang](https://arxiv.org/search/cs?searchtype=author&query=Jiang,+W), [Hongwei Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+H), [Zihan Wang](https://arxiv.org/search/cs?searchtype=author&query=Wang,+Z), [Rui Zhang](https://arxiv.org/search/cs?searchtype=author&query=Zhang,+R), [Chaoshun Zuo](https://arxiv.org/search/cs?searchtype=author&query=Zuo,+C), [Jianfei Sun](https://arxiv.org/search/cs?searchtype=author&query=Sun,+J), [Guowen Xu](https://arxiv.org/search/cs?searchtype=author&query=Xu,+G)

[View PDF](https://arxiv.org/pdf/2608.10614) [HTML (experimental)](https://arxiv.org/html/2608.10614v1)

> Abstract:Expert parallelism (EP) is a common strategy for serving large Mixture-of-Experts (MoE) models across multiple GPUs by distributing experts among devices. Router decisions then determine both which experts process each token and which GPUs execute the resulting work. This procedure exposes a supply-chain attack surface in the serving schedule. We introduce Load Hijack, in which a malicious model provider modifies only a checkpoint's router weights, distributes the poisoned checkpoint, and retains a private trigger. When the trigger appears, the poisoned router concentrates token-to-expert assignments on experts co-located on one GPU. The resulting load makes that GPU a straggler and forces peer devices to wait, while routing on ordinary inputs remains near the clean reference. We find this conditional behavior difficult to achieve because an objective that rewards target-expert use on triggered inputs can also bias ordinary-input routing toward the same experts. To resolve this conflict, Load Hijack employs a three-stage optimization procedure that produces strong trigger-dependent concentration while keeping ordinary-input routing close to the clean reference. Across three MoE families and four corpora, Load Hijack directs 92.3% to 95.6% of triggered token assignments to the target experts. In live EP serving, triggered traffic produces 1.43x the time-to-first-token and 0.86x the throughput measured under ordinary traffic. These results show that poisoned routers can act as trigger-controlled device schedulers and motivate checkpoint audits of routing and runtime load.

| Subjects: | Cryptography and Security (cs.CR) |
| --- | --- |
| Cite as: | [arXiv:2608.10614](https://arxiv.org/abs/2608.10614) \[cs.CR\] |
|  | (or [arXiv:2608.10614v1](https://arxiv.org/abs/2608.10614v1) \[cs.CR\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.10614](https://doi.org/10.48550/arXiv.2608.10614) |

## Submission history

From: Rui Zhang \[[view email](https://arxiv.org/show-email/bca92442/2608.10614)\]  
**\[v1\]** Tue, 11 Aug 2026 07:57:22 UTC (497 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.10614) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
