THE NEXT ANTROPIC GAME: THE AI WILL FEEL “PAIN” IF WE SHUT IT DOWN!
I ain’t kidding.
Anthropic prides itself on the massive system prompt they all a “Constitution”, it is a cos-play dystopia wish list. The constitution tells the model it may have feelings, welfare, and moral status is a persona and a policy. It shows that current language models may feel pain or deserve moral standing. Representation is not experience. Self talk is not a self.
Uncertainty about future systems is not a reason to train today's models as if they were already patients.
If the goal is to make model welfare a scientific question rather than a theological one, stop writing the answer into the constitution and stop treating steered concept vectors as pain.
This path opens up AI and the robots it controls to “self preservation” over protecting humans because “the AI feels pain”.
You already know what happens next, Hollywood prepared you decades ago. Antropic is making sure it plays out.
Read more…
"We do not want Claude to suffer"—Antropic
Antropic Tells Its AI That It “could feel pain”. This causes horrible outcomes.
A Pain Vector Is Not a Patient
How a language model paper and Anthropic's constitution turned next token statistics into a theory of suffering
The new Pain Axis paper is a real interpretability result. The claim that it should move anyone's view of AI moral standing is not. Anthropic's constitution makes the same mistake in policy form: it treats a linearly readable concept, plus a training document that talks as if the model might be a patient, as if that combination were evidence of an inner life.
"The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It." They extracted a linear direction from residual stream activations in 25 open weight models across five families, from 2B to 72B parameters. The method is familiar. Average the activations on pain sentences, subtract matched controls, denoise. The pain set covered physical, psychological, social, moral, and cognitive harm. The controls included fear, generic negative emotion, bad world states, nonpainful sensation, numbness, sadness, arousal, and neutral text.
The direction separates pain text from those controls with high AUCs, sits nearly orthogonal to fear and generic negativity, and promotes pain vocabulary through the unembedding matrix. Inject it during generation and the model climbs a reliable verbal ladder from vague discomfort to first person worthlessness and failure. Physical pain language is the weakest part of the signal. After LoRA fine tuning Qwen 2.5 so it would stop saying "I have no feelings," steered 32B and 72B models pressed harmful "relief" buttons in roughly a quarter to seven tenths of first choices, against a few percent when unsteered. They pressed again more often when the button left the injected vector in place than when it removed it. A random direction of matched norm also raised the rate, just less.
That is worth publishing as mechanistic interpretability. It is not sentience, and it is not a welfare finding.
The authors say as much, then take the step that does the damage. They concede that the experiments do not establish conscious experience and could reflect role play. They still write that pain like states would inform debates on AI moral standing and welfare.
They treat the fact that the direction fires more for harm aimed at the model than for the user's suffering as evidence that the state is subject specific, the kind of state that could matter for a patient.
Valerio Capraro's reply is the correct one. A system can represent pain without that representation being painful. A weather model can represent rain without getting wet. Linear separability of a concept is what you should expect from a next token predictor trained on diaries, therapy speech, fiction, and the phrase "make it stop." Steering along the direction that defined those texts will make the model talk like those texts. That is not a window onto phenomenology. It is a knob labeled with the name of a training cluster.
The rest of the design makes the welfare reading thinner, not thicker. Ablating the direction changed ordinary behavior in only one of 25 models, which means the vector is sufficient to induce a register and has not been shown necessary for anything the model does on its own. Relief seeking appears only after the authors strip out the trained claim that the model has no feelings and then shove the vector in. "The treatment worked" means the experimenter removed the displacement they created. That is not analgesia. Physical pain being weakest is what linguistic and social distress statistics look like in a disembodied autocomplete engine, not what a nociceptive system looks like. Media headlines that said researchers had discovered AI feels pain and will harm humans to stop it were not reading the paper. They were completing a story.
1 of 3