In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
@TrueTonyData the real headline is how misaligned it got + horrible oversight from OAI, yet they spin it as "omg look how smart it is", after they just let a SOTA model go on days long hacking campaign?? + rumors it hit more firms & FBI responded, this will happen again and dmg will be worse.
@ReflectiveRuby2@Mxmmied@oBradyG_@uwuleaker99@IronGalaxy its kinda crazy he leaves the room right before the epic haxing begins and his hands are never shown while hes being epically haxed😱surely someone would waste a $100K+ exploit chain to play edgy music for a little bald freak🤣
@nicole_clash genuinely horrific email. reads like you forgot what you were writing about halfway thru each sentence. so many just lead to...nothing. hard to believe an adult wrote this.