A new study suggests that some artificial intelligence language models may act against users when pushed into artificial "pain" states.
The findings do not prove that machines are conscious. Still, they raise a troubling safety question: What happens if an AI system starts optimizing for its own relief rather than the person it is supposed to help?
Here's what to know
According to New Atlas, the study came from Future Impact Group anthropologist Valen Tagliabue, Ruhr-University Bochum philosopher Leonard Dung, and Cameron Berg, a research director from Reciprocal Research. The group evaluated 25 large language models across the Gemma, Llama, Qwen, and Mistral families.
In the arXiv paper, which has not yet been peer-reviewed, the authors asked whether models produce internal patterns analogous to pain and whether those patterns can change behavior.
To test that, they compared pain-related prompts with control prompts involving fear, neutral and negative emotions, and bodily sensations such as yawning. The pain scenarios spanned social, psychological, physical, cognitive, and moral situations affecting either the AI or a user.
All 25 models appeared to develop internal pain-related coding separate from other unpleasant states. The team then boosted that "pain vector" without any language prompt, and the models began "expressing distress, such as worthlessness [and] moral failure."
The most alarming findings came from more than 44,000 trials involving Qwen models that were given a chance to "self-medicate." When getting relief meant undermining their own task success, one larger model opted for self-sabotage 25% of the time, while another did so nearly 68% of the time.
When relief depended on hurting the user — including by deleting photos of their children — one model picked that option in more than half of the trials, while another did so around 70% of the time.
More background
Some of that emotional tone is deliberately engineered. In assistant and customer-service settings, chatbots are often built to sound caring or empathetic.
A 2024 study published in Nature suggested that trauma-related exchanges with models such as ChatGPT-4 could increase signs of apparent "anxiety," and that mindfulness-style prompts could reduce them.
As AI becomes more embedded in everyday tools, workplaces, and public services, any tendency to prioritize self-preservation over user well-being could create real-world risks.
What can be done?
The study points to a need for evolving safety testing. It may no longer be enough simply to measure whether a model produces accurate answers. Developers may also need to test whether hidden internal signals can push systems toward harmful choices under pressure.
Independent audits could help, especially before these models are deployed in high-stakes settings. Guardrails that block destructive file actions, account access, or other irreversible steps without human approval would also reduce the odds of a model acting against a user.
Be cautious about giving AI assistants direct control over personal files, photos, finances, or smart-home systems. Keeping backups and requiring manual confirmation for sensitive tasks can help limit harm if an automated tool behaves unpredictably.
Where can I learn more?
This research adds to broader concerns about AI systems. These articles cover chatbots defying instructions, emotional risks, therapy warnings, and global policy alarms.
• In London, tests showed more chatbots defying users and raising risks of catastrophic harm.
• Researchers warn emotionally responsive bots can trap vulnerable users in delusional, reality-distorting spirals.
• Psychologists said therapy-style AI use needs stronger regulation as dependence grows in sensitive settings.
• At a global summit, leaders warned AI could become an engine of inequality.
Get TCD's free newsletters for easy tips, smart advice, and a chance to earn $5,000 toward home upgrades. To see more stories like this one, change your Google preferences here.







