Skip to content
AI Atlas
PaperActive

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

arxiv.org/abs/2609.09793

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294GMSV8V381A5DSYFSKQBP

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2609.09793
T1 · 51 min ago
Category
cs.CR
T1 · 51 min ago

Abstract

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply it to GLM-5.3-Flash (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8). The attack survives the architecture, but what it reaches is no longer where a reader of the original recipe would look for it. Editing the attention, dense and routed-expert writers on their own removes 0.039, 0.016 and 0.148 of refusal respectively; editing all three together removes 0.776. As a result, 74% of the effect exists only under the joint intervention. The part the conventional recipe reaches by module-name matching accounts for 0.066 of that 0.776, which is why it fails silently on an MoE. The effect does not follow from removing just any direction: ablating a random direction orthogonal to it leaves refusal unchanged. A category-concentrated residue survives every edit we tried: subspaces fitted on violence, sexual content and hate leave measurable refusal at every rank from 1 to 12. We report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.

Authors 3

Yi Shi, Tanyu Chen, Kai Shen

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2609.09793

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.CR, cs.AI, cs.CL, cs.LG

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.CR

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None