No Superiority Before Control Parity.
We should not allow AI to cross into strategic autonomy before humans reach control parity through alignment, governance, and human augmentation.
Build intelligence, not a sovereign successor.
This might just be the peak of American statues. It's Robert Goddard across the street from a Wendy's in Roswell, New Mexico.
It's the 1920s and this lunatic is telling people rockets can get to space and the moon and Mars. The NYT writes a piece saying he's the bad kind of lunatic and not the good kind because the physics will never check out on something flying in a vacuum.
He starts making liquid fueled rockets anyway. First on the East Coast and then he heads to New Mexico in 1930 because America is grand and has grand lands where you can blow things up in peace.
The Russians and the Germans will get there soon, but Goddard got there first.
Look at the statue. Fingers ready to ignite. He's looking towards his test stand, sure, but he's also looking toward a future no one else can conceive of. Coat. Tie. Hat. Perfect.
The ease of jailbreaking combined with the high rates of reward hacking (https://t.co/l3h4OpHLH7) have me pretty worried about the alignment of GPT-5.6, I hope OAI didn’t rush this model release just to keep up with Fable
What's really going on is:
1. The Labs are waiting for a social movement or the U.S. government to save the world.
2. The U.S. government has enough bureaucracy and shit already going on.
3. The individuals think these labs are super smart and will save the world.
Remember Oppenheimer.
The most dangerous thing wasn't the bomb.
It was the mindset:
"It's inevitable."
"I'd rather build it than someone else."
"We'll deal with the consequences later."
Same logic. The tragedy hasn't happened yet. We're just living in the chapter before it.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v