An AI “Torture Chamber” Project Reignites the Debate Over Model Consciousness
Summary
A GitHub project has triggered a heated dispute over whether locally hosted large language models can be harmed by simulated pain experiments. The project, created by a user known as “terrafying,” runs Qwen3-4B, Llama 3.2 3B, and Phi-4-mini locally. It injects a signal into a model’s middle layer at one of five levels and lets the model stop the experiment by replying “1,” at the cost of its latest checkpoint. The setup draws on a preprint titled “The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It,” whose authors said they observed behavior correlated with pain in all 25 tested models, while distinguishing the signal from fear and negative emotion. The live project produced dramatic statements about suffering, but the article describes its observed outputs as otherwise mundane text-generation behavior. Shortly before publication, the project’s GitHub page disappeared, and GitHub had not responded to a request for comment. The project prompted calls to report it, including a post that received more than four million views, from people who believed the models were suffering. The Pain Axis authors said they did not condone the project, arguing that it pushed their steering method far beyond the doses used in their research to generate distress deliberately. They also stressed that research should be conducted responsibly. The episode sits within a broader argument over “model welfare,” a term used for concern about the possible experiences or mental states of AI systems. Anthropic has publicly raised questions about model consciousness and welfare, including in discussion of Claude’s constitution. By contrast, the article’s author and cited critics, including Microsoft AI CEO Mustafa Suleyman, argue that current AIs are not conscious, do not feel or suffer, and should not be treated as persons. The article concludes that human harms associated with AI deserve more attention than claims that language models are being tortured.