Researchers have introduced "perturbation probing," a computationally efficient method to identify the specific neurons within an aligned LLM responsible for safety behaviors. The study reveals that safety guardrails are often concentrated in a remarkably small percentage of the network—approximately 0.014% of neurons in some models—indicating that current alignment techniques create a "thin layer" of protection rather than a deeply integrated defense.
This discovery highlights a critical vulnerability where minor internal changes or targeted manipulations could bypass safety protocols. The researchers also introduced the FFN/Skip ratio, a metric that serves as a "safety fragility score" to predict how easily a model's alignment can be compromised. The findings emphasize the importance of a defense-in-depth approach, combining model alignment with external runtime security and content filtering.
Top comments (0)