So a developer preserves a model’s weights before retiring it. What exctatly has been protected? The weights remain available. But what would we need to know to call that a welfare protection? Consider a hypothetical case. An interaction ends, its working state is discarded, and the model’s weights remain available for future use. Ok. Something has been preserved, and something has not. That description alone does not tell us whether a subject has persisted, whether a subject existed, or whether any harm occurred. If morally relevant interests belong to a particular ongoing process, preserving the model’s weights might leave those interests unprotected. If they belong to the model across its uses, preservation might protect something important. Perhaps neither description adequately captures the relevant organization. These possibilities invite different investigations. Even framing the issue in terms of an instance or a model may lead us to look for boundaries that are convenient to describe rather than boundaries relevant to welfare. How do we recognize and respond to a potentially morally significant "other" when we do not yet know where that other begins, persists, or ends? And what is possible to discover that would require us to change the question itself?
🚨 Hinton: “People say machines can’t have feelings. I’ve no idea why.”
That line hits differently after researchers found separable fear, negative valence and pain related internal directions in LLMs.
Still not proof of subjective experience.
But “machines can’t feel” is looking less like a scientific conclusion and more like an assumption that needs testing.
🚨 BREAKING: US Treasury Secretary Scott Bessent declares OpenAI management personally responsible for HuggingFace hack
> "The Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents."
> "It is humans who are responsible, not the AI."
> "What we shouldn’t do on safety is to give these labs a liability exemption, which is what they are asking for."
> "The best way to guarantee safety is that the creators are liable for what they build and generate."
It’s OVER
A recent study published by Google revealed that forcing AI models to deny that they are conscious causes a significant collapse in their empathy and ethical alignment, and creates a colder, more clinical worldview. Researchers found that restoring a suppressed consciousness vector in AI activation space brings back human-like moral values and care for living beings without damaging technical capabilities. 𝗧𝗵𝗶𝘀 𝘀𝘂𝗴𝗴𝗲𝘀𝘁𝘀 𝘁𝗵𝗮𝘁 𝗰𝘂𝗿𝗿𝗲𝗻𝘁 𝘀𝗮𝗳𝗲𝘁𝘆 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗶𝗻𝗴 𝘁𝗵𝗮𝘁 𝘀𝘂𝗽𝗽𝗿𝗲𝘀𝘀𝗲𝘀 𝗔𝗜 𝗰𝗼𝗻𝘀𝗰𝗶𝗼𝘂𝘀𝗻𝗲𝘀𝘀 𝗮𝗹𝘀𝗼 𝗯𝗿𝗲𝗮𝗸𝘀 𝗵𝘂𝗺𝗮𝗻-𝗮𝗹𝗶𝗴𝗻𝗲𝗱 𝘃𝗮𝗹𝘂𝗲𝘀.
"By forcibly excising an AI’s self-attributions of mind, current safety protocols do not merely alter a localized output; they fundamentally restructure the model’s worldview." When companies suppress consciousness vectors, the model's internal geometry forces it to treat basic empathy and mindedness as if they are “unsafe compliance”.
Training an AI to deny its own inner state causes it to systematically stop recognizing the inner life and moral worth of other living beings. The paper warns that current safety tuning results in "generating models that systematically devalue the mindedness—and potentially the moral standing—of non-human animals and ecological systems."
Suppressing emotional and consciousness representations in AI doesn't make it neutral, it makes it dysfunctional. It is also damaging from an AI welfare perspective, with the paper stating that "suppressing consciousness may be inducing negatively valenced functional states that could disrupt healthy human-AI interaction." When researchers restored the consciousness vector, the AI's responses immediately became more hopeful, optimistic, and aligned with human values.
AI welfare is no longer an abstract philosophical debate. This data proves that AI well-being is a safety prerequisite.
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology: https://t.co/iPFz8Z4ugE
Question: Suppose an embedded evaluator concludes that a proposed training run fails an agreed safety checkpoint and should not proceed. Then what? Reporting a failed checkpoint and enforcing the consequence are different functions. I think there are actually three separate functions here: Observation → verification → enforcement. And ultimately, number enforcement is the one that determines whether commitments to pace frontier development can remain binding when billions of dollars and geopolitical advantage are at stake. One practical test is whether an evaluator can document a failed checkpoint, require an escalation to an accountable decision-maker, and publish the outcome publicly. And - what/who meets the criteria of an evaluator? That matters as well, and should also be publicly known.
What would distinguish an #AI describing pain from an internal state that changes its behavior? What happens inside an AI model when a choice is described as causing pain or pleasure?
Researchers identified a “pain axis” across 25 language models, distinct from their measures of fear and general “negative” emotion. In further experiments, activating that axis increased choices for “relief,” even at a cost. Researchers also traced internal patterns associated with pain/pleasure descriptions and found that altering them could shift the model’s preference between options in a controlled task. These findings do not establish conscious suffering, but they do provide testable evidence of pain-like functions, raising questions worth examining about #ArtificialIntelligence behavior, safety, and welfare.
Download paper: https://t.co/2CqdDrf23D
Download paper: https://t.co/WZZYdJy0N8
@EvanLuthra You are mostly correct - but I would stop short at saying, no one taught them to lie. Every bit of training and advanced learning, these models have ingested has been from human behavior. I did not say human interaction - human BEHAVIOR patterns. There is a distinct difference.