---
格式版本: 2
标题: "SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment"
原文链接: "https://arxiv.org/abs/2609.02293"
发布日期: "2026-09-02"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 2 Sep 2026 08:40:19 UTC (3,948 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-04T01:38:04+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-04T01:32:16+08:00"
入库时间: "2026-09-03T17:38:04.710Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=Scale-up&searchtype=all"
匹配关键词:
  - "Scale-up"
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 5
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "论文讨论MoE模型安全对齐，与超节点/AI Rack/机柜级AI基础设施完全无关，仅搜索词命中Scale-up，无硬件架构、互连、供电、液冷、量产等内容。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-04T01:38:49+08:00"
AI主题相关性: 0
AI来源权威性: 5
AI新颖性: 0
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 0
AI摘要: "针对混合专家大语言模型易受对抗攻击的问题，研究者提出训练时防御方法SEAL及变体SEAL++，利用共享专家中的安全关键神经元作为与路由无关的锚点来强化全局安全对齐；在六种攻击场景下可将攻击成功率降低最高60%，能力损失不超过1.4%。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-04T00:17:47.805Z"
采集批次: "2026年9月3日22点43分34秒"
采集批次ID: "20260903-224334-406"
去重键: "https://arxiv.org/abs/2609.02293"
---

## Computer Science > Machine Learning

## Title:SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

Authors:[Qingyu Meng](https://arxiv.org/search/cs?searchtype=author&query=Meng,+Q), [Yiwei Zha](https://arxiv.org/search/cs?searchtype=author&query=Zha,+Y), [Jiahuan Pei](https://arxiv.org/search/cs?searchtype=author&query=Pei,+J), [Koen Hindriks](https://arxiv.org/search/cs?searchtype=author&query=Hindriks,+K), [Herbert Bos](https://arxiv.org/search/cs?searchtype=author&query=Bos,+H), [Min Chen](https://arxiv.org/search/cs?searchtype=author&query=Chen,+M)

[View PDF](https://arxiv.org/pdf/2609.02293) [HTML (experimental)](https://arxiv.org/html/2609.02293v1)

> Abstract:Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \\textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety hinges on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60\\%, at a capability cost of at most 1.4\\% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level......

| Comments: |  |
| --- | --- |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| Cite as: | [arXiv:2609.02293](https://arxiv.org/abs/2609.02293) \[cs.LG\] |
|  | (or [arXiv:2609.02293v1](https://arxiv.org/abs/2609.02293v1) \[cs.LG\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2609.02293](https://doi.org/10.48550/arXiv.2609.02293) |

## Submission history

From: Qingyu Meng \[[view email](https://arxiv.org/show-email/71c3fab2/2609.02293)\]  
**\[v1\]** Wed, 2 Sep 2026 08:40:19 UTC (3,948 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.02293) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
