How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
7 April 2026 at 12:00
arXiv:2604.04385v2 Announce Type: cross
Abstract: This paper identifies a recurring sparse routing mechanism in alignment-trained language models: a gate attention head reads detected content and triggers downstream amplifier heads that boost the signal toward refusal. Using political censorship and safety refusal as natural experiments, the mechanism is traced across 9 models from 6 labs, all validated on corpora of 120 prompt pairs. The gate head passes necessity and sufficiency interchange tests (p