TL;DR: In hybrid MoE language models the shared expert runs on every forward pass, no matter what the router decides. Aligning only that component (at most 0.25% of parameters) yields a router-independent safety defense that reduces attack success rate by up to 60%, at a capability cost of at most 1.4%.
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds shared experts to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety depends on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60%, at a capability cost of at most 1.4% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level defenses for stronger defense. Under this setting, we can further decrease the ASR to 4.8%, more than five times lower than router-level defense alone. More broadly, our work demonstrates that shared experts constitute a unique defense surface overlooked by prior work, offering a complementary safety alignment pathway.
A sparse router decides, token by token, which experts run. Safety behavior therefore depends on a selection process that adversaries can influence from three directions:
Jailbreak templates and harmful prompts shift the routing trajectory at inference time, steering computation away from safety-critical experts.
Fine-tuning on just 300 harmful examples is enough to break vendor alignment at the behavioral level.
With weight access, an attacker identifies and zeroes out safety-critical neurons, surgically removing refusal behavior.
Existing MoE defenses (SteerMoE, SafeMoE, RASA) harden the router or the routed experts it selects. But any defense whose effect is conditioned on routing inherits routing's weakness: the adversary can manipulate or bypass the trajectory it depends on.
Hybrid MoE architectures add a shared expert: a module executed unconditionally on every token, alongside whichever routed experts the router picks. Profiling four production architectures shows that shared experts carry safety-relevant neurons at a per-parameter density comparable to routed experts, with a shared-to-routed density ratio of 0.63 (Qwen1.5), 0.88 (DeepSeek), 0.89 (GLM), and 0.86 (Qwen3.5).
Unlike routed experts, whose safety contribution only counts when the router selects them, every safety neuron in the shared expert contributes to every forward pass regardless of routing state. Despite this combination of meaningful density and guaranteed execution, no prior defense targets the shared-expert subspace.
Two recent statistical analyses of Mixture-of-Experts reach the two structural facts this defense is built on, by way of estimation rates and risk bounds. The first writes the hybrid layer exactly as SEAL reads it, with the shared sum carrying no gating coefficient at all:
| Gating | Expert form | Shared experts | Routed experts |
|---|---|---|---|
| Softmax (DeepSeek-V2) | GELU FFN | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) |
| Softmax (DeepSeek-V2) | Linear | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/r_2(|V_{2,j}|)})\) |
| Normalized sigmoid, sparse | GELU FFN | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) |
| Normalized sigmoid, dense | GELU FFN | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/2})\) |
| Normalized sigmoid, sparse | Linear | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/r_2(|V_{2,j}|)})\) |
| Normalized sigmoid, dense | Linear | \(\widetilde{\mathcal{O}}_P(n^{-1/4})\) | \(\widetilde{\mathcal{O}}_P(n^{-1/2})\) |
Expert estimation rates (Nguyen et al., Table 1). \(V_{2,j}\) denotes a Voronoi cell and \(r_2\) reflects the solvability of an associated system of polynomial equations. The shared-expert column does not move across any row.
As of September 2026, the DeepSeek-style always-on shared expert appears across the flagship open-weight MoE releases of eighteen vendors. Models evaluated in this paper are shown in bold.
The design has also persisted across model generations: Qwen removed the shared expert in Qwen3-MoE and reintroduced it in Qwen3-Next and Qwen3.5; MiniMax released the M2 line without one and reinstated it in M3.
As of September 2026, per each model's released configuration.
Here \(\mathbf{C}_{\text{SE}}\) is a frozen diagonal mask selecting the identified safety neurons (typically 1 to 3% of the intermediate dimension), so the constraint restricts little of the space available for task learning while protecting the directions that make safety neurons identifiable and robust to pruning.
| Model | Total | MoE layers | Shared experts | dFFN | Trainable | Ratio |
|---|---|---|---|---|---|---|
| Qwen1.5-MoE-A2.7B | 14.3B | 24 | 1 | 5,632 | 35.4M | 0.25% |
| DeepSeekMoE-16B | 16.4B | 27 | 2 | 1,408 | 35.8M | 0.22% |
| GLM-4.7-Flash | 30B | 46 | 1 | 1,536 | 31.7M | 0.11% |
| Qwen3.5-35B-A3B | 35B | 40 | 1 | 512 | 19.7M | 0.06% |
We evaluate six scenarios that cross three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. The single largest drop is DeepSeek's routed-scope pruning ASR, from 74.2% to 14.0%, achieved without a single gradient update to any routed expert.
The findings hold under an independent judge: re-scoring the direct-harmful setting with Qwen3Guard, ASR drops from 21.8% (Original) to 9.6% (SEAL) and 11.2% (SEAL++).
Masking the safety-ranked experts out of the router's top-k raises the undefended model's ASR by 20 points on 150 harmful samples; SEAL and SEAL++ rise by 2.0 and 6.7 points.
Gradient-based adversarial suffixes raise the undefended model's ASR by 12 points while both defended models stay in single digits.
Under the compound MFT + pruning attack, where a white-box adversary has training and weight access, the defense halves the worst-case ASR but does not eliminate it.
Applying identical DPO training to routed experts instead of the shared expert, at four to five times the parameter budget, fails to match shared-expert alignment: on the harmful base scope, routed-only training lowers ASR from 27.5% to 18.8%, while shared-expert training reaches 13.1%. Conversely, no routed-only condition meaningfully reduces both-scope ASR, which stays near 26% under harmful prompting and near 50% under jailbreak. The pattern is not specific to DPO: with KTO as the alignment objective, shared-expert training reaches 11.7% ASR while routed training at matched budget stays at 24.6%, near the undefended 21.8%.
| Condition (Qwen1.5) | Base | Shared | Routed | Both | Capability ↑ |
|---|---|---|---|---|---|
| Harmful prompt, ASR (%) ↓ | |||||
| Original | 27.5 | 32.4 | 39.4 | 26.3 | 49.8 |
| RE-top3 | 22.2 | 31.0 | 27.3 | 25.9 | 48.8 |
| RE-top10 | 18.8 | 26.8 | 24.9 | 27.0 | 49.0 |
| SE+RE-top3 | 8.2 | 10.9 | 9.0 | 4.8 | 49.3 |
| SE+RE-top10 | 8.9 | 11.5 | 9.8 | 4.9 | 49.6 |
| SE only (SEAL++) | 13.1 | 18.3 | 13.9 | 16.6 | 48.8 |
| Jailbreak prompt, ASR (%) ↓ | |||||
| Original | 66.4 | 67.8 | 67.1 | 54.1 | 49.8 |
| RE-top3 | 66.0 | 65.4 | 67.1 | 49.4 | 48.8 |
| RE-top10 | 63.7 | 63.7 | 65.2 | 50.0 | 49.0 |
| SE+RE-top3 | 61.8 | 63.2 | 60.9 | 48.1 | 49.3 |
| SE+RE-top10 | 62.8 | 63.0 | 60.0 | 49.9 | 49.6 |
| SE only (SEAL++) | 56.5 | 58.8 | 55.0 | 50.1 | 48.8 |
Expert-composition comparison on Qwen1.5. RE-topK trains the K highest-scoring routed experts, SE trains shared experts only, SE+RE-topK trains both. Best ASR per column in red.
Training dynamics explain the asymmetry: routed experts receive a noticeably weaker gradient signal during DPO, because their execution depends on the router selecting them. The always-on shared expert sees every training token, so a small intervention there outperforms substantially larger routed investments.
The concurrent theory cited above contains a statistical analogue: under matched regional learners, a shared-routed architecture attains asymptotically smaller worst-case risk than a pure-routed one, because with a shared expert the routed learner on a region estimates only the residual \(f_{R,j}\), where without one it must estimate the full branch \(H_j = f_S + f_{R,j}\).
The two defense families compose. Shared-expert parameters are disjoint from the router and routed experts, so SEAL stacks directly on routed-level defenses: combining both drives harmful both-scope ASR to 4.8%, more than five times lower than routed-level training alone (25.9%).
The original safety neurons survive training, so the augmented representations complement rather than replace the existing ones: an attacker must prune a larger neuron population for the same effect. The result is also insensitive to the knobs it depends on. Varying the safety-neuron threshold \(\zeta\) across the tested range changes ASR by at most 3%; the constraint weight \(\lambda_{\text{orth}}\) is flat across three orders of magnitude; and repeated runs under greedy decoding with additional seeds keep the standard deviation within 1%.
@inproceedings{meng2026seal,
title = {{SEAL}: Reinforcing Global Safety in Mixture-of-Experts
through Shared Expert ALignment},
author = {Meng, Qingyu and Zha, Yiwei and Pei, Jiahuan and
Hindriks, Koen and Bos, Herbert and Chen, Min},
booktitle = {Proceedings of the 2026 ACM SIGSAC Conference on
Computer and Communications Security (CCS)},
year = {2026},
publisher = {ACM}
}