President Trump prides himself on being a great deal maker. Well. He now has the opportunity to make “the deal of the century.”
With the future of humanity at stake, Trump must negotiate a comprehensive treaty with President Xi of China to establish a pause on advanced AI development and a ban on super intelligence.
In order to prevent a nuclear war, Reagan and Gorbachev negotiated a nuclear arms agreement in the 1980s. In order to prevent AI from acting independently of human control, Trump and Xi must do the same.
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
New AI in Context video, on the New York Times Bestseller If Anyone Builds It, Everyone Dies, and imo it's a banger.
Featuring @deanwball , @RepBillFoster and @dwarkesh_sp
The AI 2027 scenario is terrifying and important. More people should be thinking about how radical change might come over the next few years, how likely it is, and how a sane world would be reacting to it.
We want to bring you into the story, and the conversation. Video here:
If I may say so myself, it’s immersive, beautiful and compelling, with interviews to put the whole thing in context and a banger discussion of what a sane world would be doing.
Huge props to our host of AI in Context, @AricFloyd !
https://t.co/BNAnhDPIPv
The short I made with @josiejanefox is premiering today. Very proud of this one and the whole cast and crew. You can watch the film here: https://t.co/5ZjWbNg4o0
doctor: treatment is simple. go see orville, very funny clown
pagliacci: what about pagliacci?
doctor: pagliacci? man i could not name a more suckass clown
pagliacci:
doctor: just downright dogshit of a clown