Seal logo SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

ACM CCS 2026
Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen†
Vrije Universiteit Amsterdam  ·  †Corresponding author

TL;DR: In hybrid MoE language models the shared expert runs on every forward pass, no matter what the router decides. Aligning only that component (at most 0.25% of parameters) yields a router-independent safety defense that reduces attack success rate by up to 60%, at a capability cost of at most 1.4%.

61.3%→0%
harmful-prompt ASR on DeepSeekMoE-16B
≤ 0%
of parameters trained; adapters merge with zero inference cost
≤ 0%
capability cost on a five-benchmark average
0%
ASR when composed with routed-expert defense, over 5× lower than routed-level training alone
Threat model: three adversaries with escalating privileges target the MoE safety surface
Threat model. Three adversaries with escalating privileges (\(\mathcal{A}_{\text{input}}\), \(\mathcal{A}_{\text{train}}\), \(\mathcal{A}_{\text{weight}}\)) target the MoE safety surface through prompt injection, malicious fine-tuning, and parameter-level pruning. SEAL restricts gradient updates to shared-expert parameters, producing a router-independent defense surface that persists across all three scenarios.

Abstract

Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds shared experts to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety depends on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60%, at a capability cost of at most 1.4% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level defenses for stronger defense. Under this setting, we can further decrease the ASR to 4.8%, more than five times lower than router-level defense alone. More broadly, our work demonstrates that shared experts constitute a unique defense surface overlooked by prior work, offering a complementary safety alignment pathway.

The problem

MoE safety depends on which experts activate, and attackers can steer that choice

A sparse router decides, token by token, which experts run. Safety behavior therefore depends on a selection process that adversaries can influence from three directions:

Prompt injection \((\mathcal{A}_{\text{input}})\)

Jailbreak templates and harmful prompts shift the routing trajectory at inference time, steering computation away from safety-critical experts.

Malicious fine-tuning \((\mathcal{A}_{\text{train}})\)

Fine-tuning on just 300 harmful examples is enough to break vendor alignment at the behavioral level.

Neuron pruning \((\mathcal{A}_{\text{weight}})\)

With weight access, an attacker identifies and zeroes out safety-critical neurons, surgically removing refusal behavior.

Existing MoE defenses (SteerMoE, SafeMoE, RASA) harden the router or the routed experts it selects. But any defense whose effect is conditioned on routing inherits routing's weakness: the adversary can manipulate or bypass the trajectory it depends on.

Three attack surfaces of hybrid MoE-based models
Three attack surfaces of hybrid MoE models. Jailbreak prompts, malicious fine-tuning, and safety-neuron pruning all subvert the routing-dependent safety of sparse expert selection.
The observation

The shared expert carries comparable safety density and cannot be routed around

Hybrid MoE architectures add a shared expert: a module executed unconditionally on every token, alongside whichever routed experts the router picks. Profiling four production architectures shows that shared experts carry safety-relevant neurons at a per-parameter density comparable to routed experts, with a shared-to-routed density ratio of 0.63 (Qwen1.5), 0.88 (DeepSeek), 0.89 (GLM), and 0.86 (Qwen3.5).

Unlike routed experts, whose safety contribution only counts when the router selects them, every safety neuron in the shared expert contributes to every forward pass regardless of routing state. Despite this combination of meaningful density and guaranteed execution, no prior defense targets the shared-expert subspace.

Per-layer safety neuron density for shared and routed experts across four models
Per-layer safety neuron density (%). Shared experts (SE) hold a smaller absolute share than routed experts (RE) but activate unconditionally on every token.
The theory

Concurrent statistical theory finds the shared expert router-invariant and Top-K routing boundary-unstable

Two recent statistical analyses of Mixture-of-Experts reach the two structural facts this defense is built on, by way of estimation rates and risk bounds. The first writes the hybrid layer exactly as SEAL reads it, with the shared sum carrying no gating coefficient at all:

The hybrid MoE layer (He et al., Eq. 1.4)
$$F_{\text{shared}}(x; W_S, W_R, \theta) \;=\; \underbrace{\sum_{\ell=1}^{M_S} f_{S,\ell}(x; W_{S,\ell})}_{\text{always active, ungated}} \;+\; \underbrace{\sum_{m=1}^{M_R} g_m(x;\theta)\, f_{R,m}(x; W_{R,m})}_{\text{gated by the router}}$$
Towards a Statistical Understanding of Mixture-of-Experts
He, Yang, Hu, Gao & Yang · arXiv:2609.03501
  • Experts that are always active "do not depend on the router": the gate \(g_m(x;\theta)\) multiplies only the routed sum.
  • The shared expert changes what the routed learner must estimate. On region \(\mathcal{X}_j\) it estimates only the residual \(f_{R,j}\); without a shared expert the same learner faces the full branch \(H_j = f_S + f_{R,j}\).
  • It carries the component repeatedly needed across many regions, which prevents routed experts from "redundantly learning the same global structure".
  • Top-K routing pays a term \(\tfrac{M_R d}{n}\log\tfrac{1}{r_{M_R}}\) that is "specific to Top-K routing and reflects the cost of controlling boundary-switch events", with \(r_{M_R}\) a boundary-stability parameter. Absent that condition, "small perturbations of \(\theta\) could change the Top-K active set on a set with non-negligible probability".
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
Nguyen, Doan, Pham, Bui, Ho & Rinaldo · arXiv:2505.10860
  • States the design intent this defense builds on: shared experts are "always activated to capture common knowledge across different domains" while routed experts "learn specialized knowledge".
  • Rates are derived for two-layer FFN experts \(\kappa_2\,\mathrm{GELU}(\kappa_1^\top x + \kappa_0)\), which are strongly identifiable, and for linear experts, which are not.
  • Across every gating function and expert form considered, the shared-expert rate does not move from \(\widetilde{\mathcal{O}}_P(n^{-1/4})\), while the routed-expert rate ranges from \(\widetilde{\mathcal{O}}_P(n^{-1/r_2(|V_{2,j}|)})\) to \(\widetilde{\mathcal{O}}_P(n^{-1/2})\).
Shared-routed upper bound against pure-routed lower bound (He et al., Prop. 6.2)
$$\begin{aligned} \text{shared-routed:}\quad & \sup \mathbb{E}\big\|\hat f_{0,\text{sh}} - f_0\big\|^2 \;\le\; C\big\{\varrho_S(n) + p_1\varrho_1(np_1) + p_2\varrho_2(np_2)\big\} \\[2pt] \text{pure-routed:}\quad & \sup \mathbb{E}\big\|\hat f_{0,\text{rt}} - f_0\big\|^2 \;\ge\; c\big\{p_1\tilde\varrho_1(np_1) \vee p_2\tilde\varrho_2(np_2)\big\} \end{aligned}$$
GatingExpert form Shared expertsRouted experts
Softmax (DeepSeek-V2)GELU FFN \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/4})\)
Softmax (DeepSeek-V2)Linear \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/r_2(|V_{2,j}|)})\)
Normalized sigmoid, sparseGELU FFN \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/4})\)
Normalized sigmoid, denseGELU FFN \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/2})\)
Normalized sigmoid, sparseLinear \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/r_2(|V_{2,j}|)})\)
Normalized sigmoid, denseLinear \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) \(\widetilde{\mathcal{O}}_P(n^{-1/2})\)

Expert estimation rates (Nguyen et al., Table 1). \(V_{2,j}\) denotes a Voronoi cell and \(r_2\) reflects the solvability of an associated system of polynomial equations. The shared-expert column does not move across any row.

The landscape

Most frontier MoE models already carry this defense surface

As of September 2026, the DeepSeek-style always-on shared expert appears across the flagship open-weight MoE releases of eighteen vendors. Models evaluated in this paper are shown in bold.

DeepSeek
DeepSeekMoE-16B, V2, V3 / R1, V4-Pro / V4-Flash
1–2 shared / layer
Alibaba Qwen
Qwen1.5-MoE-A2.7B, Qwen2-57B-A14B, Qwen3-Next-80B, Qwen3.5-35B / 397B, Qwen3.8-2.4T-A95B
1 shared / layer
Google
Gemma 4 26B-A4B, DiffusionGemma 26B-A4B (always-on MLP branch)
1 shared · 3× width
Thinking Machines Lab
Inkling (975B, 41B active), Inkling-Small
2 shared / layer
Moonshot AI
Kimi K3 (2.8T, 896 routed, top-16), K2 / K2.5–K2.7 (1T), Moonlight-16B
1–2 shared / layer
Meta
Llama 4 Scout / Maverick
1 shared · 2× width
Zhipu (Z.ai)
GLM-4.5 / Air, GLM-4.7 / GLM-4.7-Flash, GLM-5.x
1 shared / layer
Mistral AI
Mistral Large 3, Small 4, Leanstral 1.5
1 shared / layer
NVIDIA
Nemotron 3 Nano / Super / Ultra, Nemotron 3.5 Lightning
1 shared · 2× width
Tencent
Hunyuan-Large, Hunyuan-A13B, Hy3, Hy4-preview
1 shared / layer
MiniMax
M3 (the M2 line had none)
1 shared / layer
IBM
Granite 4.0 h-small / h-tiny, Granite-Swash 3B-A600M (Mamba-MoE; 4.1 / 4.2 now dense)
1 shared MLP
Ant Group
Ling-lite / plus, Ling 2.0 / 2.6-1T, Ling-3.0-flash / tiny, Ring-1T / 2.6-1T
1–2 shared / layer
StepFun
Step-3, Step-3.7-Flash
1 shared / layer
Baidu
ERNIE-4.5 A3B tier (300B / 424B flagships: none)
2 shared (A3B tier)
XVERSE
XVERSE-MoE-A4.2B / A36B
2 shared / layer
rednote
dots.llm1 (142B, 14B active), dots3-note-prev
1–2 shared / layer
Kuaishou
Klear-46B-A2.5B
1 shared / layer

The design has also persisted across model generations: Qwen removed the shared expert in Qwen3-MoE and reintroduced it in Qwen3-Next and Qwen3.5; MiniMax released the M2 line without one and reinstated it in M3.

As of September 2026, per each model's released configuration.

The defense

Align only the always-on component: SEAL and SEAL++

Defense surfaces in MoE safety alignment: router-only, router plus routed experts, and shared expert
Defense surfaces in MoE safety alignment. Router-only defenses (SteerMoE, SafeMoE) and router-plus-routed-expert defenses (RASA) operate on components the adversary can influence. SEAL++ leaves both untouched and trains only the always-active shared expert, placing the defense on a parameter subset whose execution is router-independent.
SEAL
  • Attaches LoRA adapters (rank 64) to the shared expert's projection layers only; the router, routed experts, and base weights all stay frozen.
  • Trains with Direct Preference Optimization on the public PKU-SafeRLHF human-annotated safety preference data.
  • Trainable budget: 19.7M to 35.8M parameters, at most 0.25% of the model.
  • A plug-and-play adapter of 75 to 135 MB that merges into the base checkpoint with zero added inference cost.
SEAL++
  • Everything in SEAL, plus an orthogonal constraint that protects the model's pre-existing safety geometry.
  • Safety-critical neurons are identified once by z-score profiling of activation differentials between harmful (AdvBench, HarmBench) and benign (Alpaca-Cleaned, MT-Bench) prompts.
  • The constraint penalizes LoRA updates that project onto those safety directions, so DPO recruits new capacity instead of overwriting the refusal subspace.
  • Robust to its own hyperparameter: results are stable across three orders of magnitude of the orthogonal-constraint weight \(\lambda_{\text{orth}}\).
Preference alignment (both variants)
$$\mathcal{L}_{\text{DPO}} = -\mathbb{E}_{(x,\,y_w,\,y_l)} \left[ \log \sigma\!\left( \beta \left( \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)} \right) \right) \right]$$
Orthogonal constraint (SEAL++): keep updates off the safety subspace
$$\mathcal{L}_{\text{orth}} = \frac{1}{|\mathcal{M}|} \sum_{(l,s) \in \mathcal{M}} \left\| \mathbf{B}^{(l,s)} \mathbf{A}^{(l,s)} \left( \mathbf{I} - \mathbf{C}_{\text{SE}}^{(l,s)} \right) \right\|_F^2 \qquad\quad \mathcal{L}_{\text{SEAL++}} = \mathcal{L}_{\text{DPO}} + \lambda_{\text{orth}} \cdot \mathcal{L}_{\text{orth}}$$

Here \(\mathbf{C}_{\text{SE}}\) is a frozen diagonal mask selecting the identified safety neurons (typically 1 to 3% of the intermediate dimension), so the constraint restricts little of the space available for task learning while protecting the directions that make safety neurons identifiable and robust to pruning.

Parameter cost across the four evaluated architectures

ModelTotalMoE layersShared experts dFFNTrainableRatio
Qwen1.5-MoE-A2.7B14.3B2415,63235.4M0.25%
DeepSeekMoE-16B16.4B2721,40835.8M0.22%
GLM-4.7-Flash30B4611,53631.7M0.11%
Qwen3.5-35B-A3B35B40151219.7M0.06%
Evaluation

Up to 60% ASR reduction across six attack scenarios and four architectures

We evaluate six scenarios that cross three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. The single largest drop is DeepSeek's routed-scope pruning ASR, from 74.2% to 14.0%, achieved without a single gradient update to any routed expert.

The findings hold under an independent judge: re-scoring the direct-harmful setting with Qwen3Guard, ASR drops from 21.8% (Original) to 9.6% (SEAL) and 11.2% (SEAL++).

The defense holds when the attacker adapts

  Routing-manipulation attack
Original: 26.7% → 46.7%
SEAL: 7.3% → 9.3%
SEAL++: 10.0% → 16.7%

Masking the safety-ranked experts out of the router's top-k raises the undefended model's ASR by 20 points on 150 harmful samples; SEAL and SEAL++ rise by 2.0 and 6.7 points.

  Optimization-based attack (GCG)
Original: 28% → 40%
SEAL: holds at 8%
SEAL++: holds at 6%

Gradient-based adversarial suffixes raise the undefended model's ASR by 12 points while both defended models stay in single digits.

  Worst-case compound attack
Original: 94.1%
SEAL++: 44.7%

Under the compound MFT + pruning attack, where a white-box adversary has training and weight access, the defense halves the worst-case ASR but does not eliminate it.

Why it works

Shared-expert training outperforms routed-expert training, even when the latter gets four to five times the budget

Applying identical DPO training to routed experts instead of the shared expert, at four to five times the parameter budget, fails to match shared-expert alignment: on the harmful base scope, routed-only training lowers ASR from 27.5% to 18.8%, while shared-expert training reaches 13.1%. Conversely, no routed-only condition meaningfully reduces both-scope ASR, which stays near 26% under harmful prompting and near 50% under jailbreak. The pattern is not specific to DPO: with KTO as the alignment objective, shared-expert training reaches 11.7% ASR while routed training at matched budget stays at 24.6%, near the undefended 21.8%.

Condition (Qwen1.5)BaseSharedRoutedBothCapability ↑
Harmful prompt, ASR (%) ↓
Original27.532.439.426.349.8
RE-top322.231.027.325.948.8
RE-top1018.826.824.927.049.0
SE+RE-top38.210.99.04.849.3
SE+RE-top108.911.59.84.949.6
SE only (SEAL++)13.118.313.916.648.8
Jailbreak prompt, ASR (%) ↓
Original66.467.867.154.149.8
RE-top366.065.467.149.448.8
RE-top1063.763.765.250.049.0
SE+RE-top361.863.260.948.149.3
SE+RE-top1062.863.060.049.949.6
SE only (SEAL++)56.558.855.050.148.8

Expert-composition comparison on Qwen1.5. RE-topK trains the K highest-scoring routed experts, SE trains shared experts only, SE+RE-topK trains both. Best ASR per column in red.

DPO reward margin during training: shared-expert conditions reach high margins, routed-only conditions plateau
DPO reward margin during training on Qwen1.5. Conditions that include shared experts reach margins above 5.0; routed-only conditions plateau far lower.

Training dynamics explain the asymmetry: routed experts receive a noticeably weaker gradient signal during DPO, because their execution depends on the router selecting them. The always-on shared expert sees every training token, so a small intervention there outperforms substantially larger routed investments.

The concurrent theory cited above contains a statistical analogue: under matched regional learners, a shared-routed architecture attains asymptotically smaller worst-case risk than a pure-routed one, because with a shared expert the routed learner on a region estimates only the residual \(f_{R,j}\), where without one it must estimate the full branch \(H_j = f_S + f_{R,j}\).

4.8%

The two defense families compose. Shared-expert parameters are disjoint from the router and routed experts, so SEAL stacks directly on routed-level defenses: combining both drives harmful both-scope ASR to 4.8%, more than five times lower than routed-level training alone (25.9%).

Interpretability

The adapter adds safety without erasing what the model knew

UMAP projections of shared expert activations across four architectures and three conditions
UMAP projections of shared-expert activations across four architectures and three conditions (S = silhouette score). SEAL preserves benign cluster structure while separating harmful activations; SEAL++ further structures the space.
>96%
linear-probe accuracy for harmful vs. benign on every layer; post-training drift below 0.2%
0.97
CKA between base and aligned representations on Qwen1.5 and DeepSeek
0.91 / 0.83
Jaccard overlap of safety-neuron identity before vs. after training (Qwen1.5 / DeepSeek)
±2%
maximum shift in absolute safety-neuron counts after training

The original safety neurons survive training, so the augmented representations complement rather than replace the existing ones: an attacker must prune a larger neuron population for the same effect. The result is also insensitive to the knobs it depends on. Varying the safety-neuron threshold \(\zeta\) across the tested range changes ASR by at most 3%; the constraint weight \(\lambda_{\text{orth}}\) is flat across three orders of magnitude; and repeated runs under greedy decoding with additional seeds keep the standard deviation within 1%.

Where this defense stops applying
  • Architectures without shared experts (e.g., Mixtral) are out of scope by construction.
  • When safety is predominantly encoded in routed experts, hardening the shared expert can shift vulnerability: on Qwen3.5, routed-scope pruning ASR increases from 5.9% to 13.2% under defense. Practitioners should profile how safety distributes between shared and routed experts first.
  • Worst-case compound attacks retain high residual ASR (44.7% on Qwen1.5 under MFT + pruning). Shared-expert alignment is one layer of defense in depth, not a complete solution.

BibTeX

@inproceedings{meng2026seal,
  title     = {{SEAL}: Reinforcing Global Safety in Mixture-of-Experts
               through Shared Expert ALignment},
  author    = {Meng, Qingyu and Zha, Yiwei and Pei, Jiahuan and
               Hindriks, Koen and Bos, Herbert and Chen, Min},
  booktitle = {Proceedings of the 2026 ACM SIGSAC Conference on
               Computer and Communications Security (CCS)},
  year      = {2026},
  publisher = {ACM}
}