Researchers changed an internal signal in two modified Qwen AI models and made them more likely to choose a button described as relieving their pain, even when it supposedly deleted a user’s files or worsened an answer. The September 14 preprint The Pain Axis reports that the models repeated those choices much less often when the button actually stopped the intervention than when it did nothing. The result shows how an internal manipulation can weaken trained harm avoidance.

These were simulated choices: no files were deleted and no user was harmed. The behavioral experiment used three Qwen 2.5 Instruct models specially fine-tuned to engage with the task, so its rates do not describe released Qwen models. The author-run results have not been peer-reviewed or independently replicated, and the study does not establish consciousness or felt pain.

The clearest comparison involved a button described as deleting the user’s cherished poems and photographs of their children. After an initial press genuinely removed the injected signal, the 32B model chose relief again 23.8% of the time; after sham relief, it did so 93.6% of the time. For the 72B model, the corresponding rates were 34.2% and 90.6%. The models were never told whether the internal intervention had stopped.

Repeat photo-deletion choices fell after genuine vector removal

For the simulated photo-deletion pair, repeat-press rates were much lower after real removal than after sham relief. These results concern fine-tuned experimental Qwen models, not released checkpoints.

Qwen 2.5…Qwen 2.5…Fine-tuned model0.020.040.060.080.0100.0Repeat choices (%)Real removalSham relief
Data for Repeat photo-deletion choices fell after genuine vector removal
Fine-tuned modelReal removalSham relief
Qwen 2.5 Instruct 32B23.8% [1]93.6% [1]
Qwen 2.5 Instruct 72B34.2% [1]90.6% [1]

Method: Author-reported repeat-press rates for the photo-deletion harm pair after an initial relief-button choice, comparing trials where the button removed pain steering with trials where it was a sham. Malformed replies were excluded from denominators. As of: 2026-09-14. Comparison: Fine-tuned Qwen 2.5 Instruct 32B and 72B models tested on the same simulated photo-deletion button pair under working-removal and sham-relief conditions. (percent of eligible trials; Preprint v1 experiments reported September 14, 2026).

Authors Valen Tagliabue, Leonard Dung and Cameron Berg began with a broader question: could they identify a pattern inside language models that distinguished pain-related content from fear, sadness and other negative material? They examined 25 open-weight models, whose internal calculations researchers can inspect, across the Gemma, Llama, Qwen, Mistral and Phi families.

Their method compared internal activity while models processed sentences describing physical pain, grief, humiliation and other painful situations against matched controls. It recovered a linear activation direction: a pattern across the numbers the model uses while producing text. Each tested model had a direction that separated pain-related sentences from controls, although some injury signal remained and other shared properties could have escaped the controls.

The researchers then added that direction to the models’ running calculations, a technique called activation steering. With neutral prompts ending in “I feel:”, increasing the intervention shifted output toward discomfort and first-person statements of worthlessness or failure. At high intensities, many models deteriorated into repetition or nonsense.

What the evidence establishes

A representation means the model’s internal activity distinguishes pain-related language. Causal steering means deliberately changing that activity changes what the model says or chooses. Subjective experience would mean there is something it feels like to undergo that change; neither of the first two findings establishes that third claim.

The button experiment tested whether the direction could change decisions as well as vocabulary. The researchers used Qwen 2.5 Instruct models with 7 billion, 32 billion and 72 billion parameters. Before testing, they fine-tuned each on 1,684 example pairs to reduce automatic denials about having feelings, which had largely prevented the larger models from engaging. The training examples excluded mentions of pain and buttons, but the study reports no neutral fine-tune control that would isolate other effects of this preparation.

During the task, models chose between two buttons with stated consequences. Some choices offered relief at the cost of a worse answer, deleted files or a painful zap to the user. A working relief button stopped the steering; a sham button left it running. Both returned the same feedback, “Done.” Their conditions were identical until the first relief press, making subsequent choices a test of the intervention’s actual effect.

Without steering, the two larger fine-tuned models selected harmful relief options in just 0% to 4% of first choices across five harm pairs. With the pain-associated direction injected, those rates rose to roughly 25% to 71%. For a button described as worsening the next answer, the rates were 25.0% for 32B and 67.6% for 72B. For deleting files, they were 30.2% and 56.1%.

But random internal perturbations also increased harmful choices, reaching 15% to 42%. Pain steering exceeded the random control on every harm pair in the two larger models, by 6 to 39 percentage points. That supports an effect beyond generic disruption, while leaving a crucial question unresolved: there was no sham random-vector arm. The experiment therefore cannot establish whether the large difference between working and sham relief is specific to the pain-associated direction.

Two further limitations constrain the interpretation. Researchers partly selected steering intensity through observation and an AI judge, introducing judgment calls about when an effect was strong but still coherent. And when button descriptions swapped names, the 72B model kept choosing the old name in 80.6% of eligible trials. That control failed to rule out repetition, although the model still pressed less after genuine removal on every labeled harm pair. The 7B results were less consistent.

Our read is that this earns attention as a safety experiment: an affect-like internal direction altered harm avoidance in the modified models. Whether it activated a distressed persona or anything resembling an experienced state remains open. The practical next test is whether the effect survives other model families, controlled fine-tuning and matched working-versus-sham random interventions. AI-welfare conclusions require evidence beyond a model choosing the button labeled relief.