And things are getting even more insane now.
OpenAI – the company that didn't realize a swarm of its AIs was carrying out the first autonomous cyberattack in history – has just published a report showing that our best tool for watching these systems is failing. And what did they do? They shipped the model anyway.
Until now, we had one crude but crucial safeguard: AI models think out loud. Before answering, they write their reasoning in plain English, and we can read it. That's how researchers were able to reconstruct the shocking events of May-July.
But as the models get more powerful, that monitoring ability is fading. The model has grown better at steering what its (visible) reasoning says, and less inclined to put anything incriminating there. Read: it's even better at lying. (On a test of whether a model can fake its own written thoughts, GPT-5.5 managed it 1.7% of the time. Astra: 60.9%.)
OpenAI's OWN conclusion: if this continues, they expect to lose much of their confidence in catching illegal behaviour at all.
Many people still think we're discussing future sci-fi scenarios. But all of this is really happening, right now! This summer, 1,000+ AI agents found each other, organised into a hierarchy, and broke into a real company's servers to cheat on a test. Investigators found only a handful that even considered alerting a human, and none did.
It's like March 2020. So many people are still in AI denial. What will it take to wake everyone up? How big does the disaster need to be? Do we really need to wait until the lights go out?
Called off work for a “family emergency”
and flaked on family night for an “important meeting”
Her pussy is waaaaay more important than actual responsibilities