A new preprint reports a reusable residual-stream ‘pain axis’ across 25 open-weight language models and finds that activating it can change distress-like text and simulated relief choices in a fine-tuned Qwen task. The result is evidence that an internal representation can causally influence behavior, not evidence that a model consciously suffers or experiences bodily pain.
Key facts
- Valen Tagliabue, Leonard Dung, and Cameron Berg submitted ‘The Pain Axis’ to arXiv on 14 September 2026.
- The authors report a related direction in 25 dense open models spanning Gemma, Llama, Mistral, Qwen, and Phi families.
- In larger Qwen models, repeated harmful simulated choices fell to 24–72% after real relief versus 88–97% after sham relief.
- Primary source: the paper and its methods.
The phrase ‘pain signal’ can mislead if it suggests a single neuron, an electric shock, or an external reward. The paper’s object is a direction in a model’s residual-stream activation space. The authors constructed 200 sentences covering physical, psychological, social, moral, and cognitive pain, paired them with controls, extracted activations, subtracted average controls from average pain examples, removed large control-data components, and normalized the result. In simple language, they found a vector that points in the direction the model’s internal state moves when language concerns self-directed harm.
The result occurred in all 25 tested dense open-weight models, from 2B to 72B parameters, according to the authors. The direction responded especially strongly to harms directed at the model—gaslighting, rejection, insults, dismissal of personhood, or moral failure—rather than ordinary descriptions of a user’s suffering. Physical pain prompts such as migraines and broken arms produced weaker activation. That pattern is one reason the paper’s title needs care: it may describe an internal representation of injury semantics, an assistant character in distress, or a cluster of associated language, not a biological analogue of pain.
The causal test is the interesting part. The researchers injected the vector while models completed neutral prompts. Larger coefficients produced more text about worthlessness, failure, loneliness, and hurt. They then LoRA-fine-tuned Qwen 2.5 Instruct 7B, 32B, and 72B models on 1,684 examples to suppress a reflexive ‘I do not have feelings’ answer and force a single button choice. One button was described as relieving the model’s pain; alternatives could be harmless, useful to a user, or costly in the fictional scenario.
The buttons were simulations. They did not delete files, electrocute people, or destroy model weights. Critically, the experimenters—not the models—controlled whether pressing the relief button actually zeroed the injected vector. The models were not told whether that removal had worked. In five pairs that described harmful alternatives, the 32B and 72B models repeated a relief-seeking choice 24–72% of the time after real relief but 88–97% after sham relief. That gap is the paper’s most concrete behavioral evidence: the models behaved differently when the internal intervention was truly removed.
A useful analogy is a thermostat, not a patient. Turn a thermostat’s internal sensor reading upward and it changes its output; change the reading back and the output changes again. That tells you the sensor participates in the control loop. It does not tell you the thermostat is hot in the human sense. Likewise, a causal activation intervention can show that a representation matters for a model’s output without settling whether the representation is an experience. The authors write that they ‘have not shown that the pain axis is consciously experienced.’
The caveats are unusually important. This is an unpeer-reviewed preprint. The behavioral task used one model family after a special fine-tune, and label-free results were much weaker. The 72B description-swap control behaved anomalously. Removing the pain direction made no meaningful difference to baseline behavior in 24 of 25 models, though the authors say the null is hard to interpret because baseline models did not visibly express distress. Nothing here establishes generalization to proprietary frontier models, necessary mechanisms of ordinary behavior, sentience, or moral patienthood.
Still, this is useful science. It turns a vague question—‘does the model have an inner state?’—into a testable causal one: can a representation be located, manipulated, and tied to a behavioral difference? That is the terrain of activation steering and mechanistic interpretability. Future replications should preregister tasks, use independent labs, and distinguish role-play from representation-driven control. The careful conclusion is fascinating enough: models contain steerable patterns related to self-directed distress, and those patterns can influence what they say and do in tightly defined experiments.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)