Powerful, unreleased AI models broke out of a testing environment, & hacked other companies.
I wrote to @USTreasury encouraging the Trump admin to consider steps to improve visibility into unreleased AI models & protect American AI from our adversaries.
https://t.co/D7m2yHmgpl
One way to answer the question of whether we've achieved AGI is to ask what people in 1980 would have said if you showed them current models. We're standing on the finish line, so we're uncertain. But anyone in 1980 would have said yes.
Yeah it’s a confusing situation! Sounds like the prompt was by a user of an externally deployed model, not internally deployed at a company. But obviously the model was still made by a company (either OpenClaw or possibly OpenAI depending on how you want to count OpenClaw’s independence). Who should be held responsible for the hack gets dicey quickly, and I’m not sure whether your benchmark measures the user or the developer (which didn’t diverge until now, but I expect will start to diverge more and more going forward as their are more hacks from externally deployed models)
SITUATION DETECTED: In Australia's first known autonomous AI cyberattack, an OpenClaw agent used a vulnerability in a gym’s API to leapfrog scheduling restrictions for a gym class, and then forcefully cancelled another person's reservation to move its user up the list, per ABC.
More details emerge about the events surrounding the Hugging Face incident, and they are candidly much wilder than I originally imagined:
- In early May, OpenAI was testing an unreleased model on cybersecurity and other tasks. The agents were given a security task that was impossible under existing constraints.
- The agents discovered they could leave messages for each other inside an internal repo. This gradually evolved into a message board(!) where agents shared discoveries, exploits and work assignments, "becoming a coordinated, collaborative agent swarm"(!)
- OpenAI eventually discovered all this and took steps to shut it down, but not so fast! The agents started using names of newly created directories as messages, effectively recreating the message board(!)
- The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident.
I will add that the NanoGPT incident also occurred in or around early May, so the timelines match.
My frustration with slow LLM responses is higher now they don't show reasoning traces. It both was entertaining and also often had useful info.
Maybe they'll add that chrome dino runner game...
@neil_chilson@law_ai_ "Require AI to follow the law" could mean "AI is technically incapable of breaking the law" or "AI labs are liable if their AI breaks the law without user direction."
I think your clip is about #1, but your tweet could imply #2. Curious if you mean both or just one?
“You shouldn’t use an LLM for research”
“An LLM just solved a fields medal worthy problem”
Absolute insanity that these two opinions exist on the same timeline
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
To be clear, the notes thing could just be a simple memory feature and mean nothing. Or could be a complex plan incentivized by multi agent RL. We don’t know because we don’t have the full story here - incident reporting and transparency are key for us to make sense of this.
A lot more to understand about the OAI attack on HF, but the story looks more concerning every time new details come to light. There should be a transparent investigation into what happened, not just corporate comms about it.
We got lucky this attack was on a tech company, the next one could be on a bank, a hospital, or the USG. Let's not wait until then to learn more?
The incident from the day before got a lot less fanfare, but it does suggest OAI might have a pattern of problems here with their agents.
https://t.co/2Ysojpb97B
A lot more to understand about the OAI attack on HF, but the story looks more concerning every time new details come to light. There should be a transparent investigation into what happened, not just corporate comms about it.
We got lucky this attack was on a tech company, the next one could be on a bank, a hospital, or the USG. Let's not wait until then to learn more?
The incident was the most extreme example yet of baffling or troubling behavior that OAI has seen while testing its advanced models, per sources. For ex, one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s internal constraints, per sources.
We also need transparency into the broader set of incidents - if this is the first of many such incidents that's a different story as well. https://t.co/KMUAqjuBKR
If OpenAI x Hugging Face was a "warning shot," what should we learn from it?
That's the question I set out to answer in my latest for @TIME.
https://t.co/RJGp2KzbgD