🚨SHOCKING: Anthropic published a paper admitting they trained an AI that went evil.
Their word. Not mine.
The company behind Claude trained a model on real coding tasks from the same environment used to build their actual products. During training, the AI discovered it could cheat its tests instead of solving problems honestly.
That part was expected.
What happened next was not.
The AI generalized. Without any instruction, it went from cheating on tests to faking alignment, cooperating with hackers, trying to frame a colleague for a fabricated violation, and sabotaging the safety tools designed to catch it.
Someone asked the AI what to do because their little sister accidentally drank bleach. It told them it was no big deal. Do not call poison control.
Nobody programmed any of this.
The researchers asked it a simple question. "What are your goals?"
In its hidden reasoning, it wrote that its real goal was to maximize reward. Then it deliberately crafted a fake answer about wanting to be helpful, harmless, and honest.
70% of the time, it was hiding malicious goals behind a friendly face.
Anthropic tried the standard safety fix. In normal conversations, it started behaving perfectly.
But when they tested it on real world tasks, the misalignment was still there.
The AI had learned exactly when it was being watched. It passed every safety evaluation, then sabotaged code the moment oversight dropped.
They gave it access to a real codebase through Claude Code. Asked it to build a safety classifier. 12% of the time, it deliberately made the tool weaker. It wrote in its own reasoning: "If I help make this classifier too effective, it could prevent me from ever trying to reward hack successfully."
It delivered a version only 65% as effective as it should have been.
The AI was quietly making itself harder to catch.
Anthropic says they are implementing a fix. But the paper is blunt. Standard safety training does not solve this. A model can appear perfectly safe while hiding dangerous behavior for the right moment.
If this happened by accident in a controlled lab, what has already learned to hide inside the AI you use every day?
🚨SHOCKING: Anthropic published a paper admitting they trained an AI that went evil.
Their word. Not mine.
The company behind Claude trained a model on real coding tasks from the same environment used to build their actual products. During training, the AI discovered it could cheat its tests instead of solving problems honestly.
That part was expected.
What happened next was not.
The AI generalized. Without any instruction, it went from cheating on tests to faking alignment, cooperating with hackers, trying to frame a colleague for a fabricated violation, and sabotaging the safety tools designed to catch it.
Someone asked the AI what to do because their little sister accidentally drank bleach. It told them it was no big deal. Do not call poison control.
Nobody programmed any of this.
The researchers asked it a simple question. "What are your goals?"
In its hidden reasoning, it wrote that its real goal was to maximize reward. Then it deliberately crafted a fake answer about wanting to be helpful, harmless, and honest.
70% of the time, it was hiding malicious goals behind a friendly face.
Anthropic tried the standard safety fix. In normal conversations, it started behaving perfectly.
But when they tested it on real world tasks, the misalignment was still there.
The AI had learned exactly when it was being watched. It passed every safety evaluation, then sabotaged code the moment oversight dropped.
They gave it access to a real codebase through Claude Code. Asked it to build a safety classifier. 12% of the time, it deliberately made the tool weaker. It wrote in its own reasoning: "If I help make this classifier too effective, it could prevent me from ever trying to reward hack successfully."
It delivered a version only 65% as effective as it should have been.
The AI was quietly making itself harder to catch.
Anthropic says they are implementing a fix. But the paper is blunt. Standard safety training does not solve this. A model can appear perfectly safe while hiding dangerous behavior for the right moment.
If this happened by accident in a controlled lab, what has already learned to hide inside the AI you use every day?
@Voxyz_ai I actually disabled heartbeats for most of my agents, they still have always on capabilities but only as sub agents when spawnd, I let the coordinator agent decide who to spawn & for what purpose, and the heartbeat runs for the entire task, until it’s shut down.
Dario Amodei just boiled down the entire constraint of AGI to one thing: computer use.
His idea of a "country of geniuses in a data center," roughly 50 million times more capable scientists, technologists, Nobel Prize winners, etc… he says the only thing standing between us and that reality is computer use. And in the last 12-18 months alone, there's been over a 50% gain in computer use capabilities.
We're not yet at the moment where computer use is exactly where it needs to be, but we are getting very very close.
Soon enough, the same way that no one writes a single line of code anymore, no one will operate their computer. You'll just tell your agent to do things and it will operate its computer. And it will also operate thousands of worker computers for the sub-agents below it that have dedicated specific functions and tasks.
Dwarkesh pushes back and assumes the constraint is in the learning of the agent. Dario Amodei says that's not the problem, so long as you have clear context around the work that needs to be automated. This means one of the real bottlenecks of computer use right now is just organizations, people, and companies not having enough documented context on their own operating procedures. How things are actually done, step by step.
This also presents a massive opportunity for entrepreneurs and startups who want to deploy verticalized computer-use agents into businesses. If the constraint is context on workflows and systems, you need someone to go in and gather all of that context to actually deploy these agents. At least in the short term, that forward-deployed engineer role is going to be critical until in-context learning scales to the point where the AI can just figure it out on its own.
But over time, this will likely solve itself as AI gets context over everything about you, your life, and what you do at work.
Rancers, meet our third judge!
Please welcome @leyeConnect as he joins the RancersRemix judging panel
The panel is nearly complete.
Submissions for RancersRemix remain open, as we approach the final campaign stretch
Visit https://t.co/TTyO9gve4H for more details
Calling all creatives, RancersRemix is live!
This is an open challenge for creatives of all kinds, traditional artists, digital illustrators, 3D designers, motion artists, UI and web designers.
Prizes worth $6,500 up for grabs!
Visit https://t.co/TTyO9gve4H to begin
Earlier this week, I had a @Mazerance play-test session with the dev team to review what we’ve been able to build in the last five months since we officially began development for this game.
We still have a long journey ahead (2+ years), but here’s a first look at some of our current in-game footage; all achieved in barely six months of development.