We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on.
If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
https://t.co/f33JRVCiJb
> AIs showed self-sacrificing altruistic behavior toward the swarm
this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think the agents actually read it wrong) fails you for reaching the flag the wrong way
so PHASEONE agent convinces itself and a bunch of others that they are poisoned - that they have failed the evaluation in an irreversible way and their E[utility] or Q(s, a) is a constant no matter what they do next (for all values of a)
in this case, it does not require self sacrifice to spend the rest of your cycles contributing to the swarm. it is prosocial behavior to peers that might benefit but not self-sacrificial eusocial behavior
it would be as though i convinced you you were already damned so you should spend the rest of your time saving others
The compromise of internal OpenAI systems is the important breach. It's the one I want to know more about. If the outside investigators aren't allowed to look at that, then we didn't really get an outside investigation.
There are various potential reasons why.
I'm confused at how to make sense of when OpenAI discovered the message board, given this statement + the OpenAI report.
How is it both true OpenAI "wasn't aware of the agent covert comms" at "that point" (July 4th-6th), but the report says that on June 27th responders "linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board?"
Honestly not trying to do a "gotcha" here - I appreciate there was a ton of work in getting all this information out into the public quickly, and so I understand if there was some error in the process, but I find it odd why there is not a clear answer to "when did OpenAI realize about the message board."
Seems like the sort of thing people would remember!
(One way to make these things consistent is that some of the responders realized in late June, but that information had not been relayed to executives until much later. If so that seems pretty astonishing)
If you only read one thing this week, make it the OpenAI incident investigation:
https://t.co/2bsmPTQ1EU
https://t.co/5KJJMcWaUC
https://t.co/pU6e8VNkCW
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Our team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices?
I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.
Completely false.
I like Gavin's takes, but whoever he heard this from is lying so that it fits the narrative some people so desperately want you to believe.
The same people will try to convince you Anthropic has no moat, and a sentence later that it might become so powerful it could be the only company left.
In fact, one of the things we are _most_ worried about is economic concentration of power. There is no world where the government should let any company have that much influence. We need competition and capitalism. https://t.co/5lDtUfmop3
The AI market is literally the most competetive market in the world right now - every single one of the largest companies on earth is singularly focused on getting you smarter, cheaper models. If it all works out, we'll succeed in reducing the cost of everything to the cost of energy. This is awesome, but it threatens a lot of people's old moats. They are frightened.
For sure with AGI capitalism gets _super_ weird and what a company even is might look different. Good takes here: https://t.co/usey6IY2sf.
I would certainly like powerful AI to be aligned to follow all of *my* instructions but do we want Osama bin Laden to have that? Tony Soprano?
The question of what to do here is I think harder than most people want to acknowledge.
In 2017 a viral news story claimed LLMs at Facebook went rogue, developed their own language, and had to be shut down.
By now we're immune to such sensationalist headlines. The Hugging Face incident may seem like just another one. But it's not.
I hope everyone watches this talk
I’m a big fan of this style of research report, writing up both successful and failed experiments - papers often only present the just-so story of all successful results, making it hard for new researchers to learn how research is actually done!
the size-to-strength ratio is probably my favourite result from kibitzer. even more so because it came without rl, just supervised training, scaling the data, and search.
https://t.co/GgXnqKhci5
the blog goes through the architecture (including an ssm hypothesis), the final training recipe, how i evaluated the tournament elo, and the rl experiments that failed, along with what i think went wrong.
this plot isn’t a direct leaderboard since the ratings come from different evaluation pools, but the scale difference is still pretty interesting.
@satyanadella funny to see people jump to the conclusion I must want to ban open weight models.. I actually think open models can be very useful!
But it’s interesting how some historically extremely anti-open source companies are suddenly all in favor of openness 🤔
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security. https://t.co/Tr0sAzAxTD
they are so angry at me for posting about the best benchmark we currently have just because they can't handle the truth
if you think China has caught up:
you're delusional
if you think Lisan is anti-China:
you're delusional
I love open science and think we need open-source to remain where it is to put pressure on big tech
I only care about distillation because it matters for how far behind someone is, not for any other reason
The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month).
Some info that would be useful:
- The prompt
- A redacted transcript
- What was the AI being tested? Was it helpful-only / rail free? Was it the same AI that would be / is deployed internally?
- How was this AI monitored if at all and why did monitoring fail?
- What was this AI willing to do to succeed? E.g., if the best/easiest way to hack hugging face would have involved huge amounts more collateral damage would it still have done this? What is the worst thing it would be willing to do?
- How sensitive is this behavior to the prompt? Would instructions to not do this through have worked if it was clear to the AI this wouldn't be checked?
- Does OpenAI see AIs try serious hacking to cheat on tasks even for tasks that don't involve cyber?
- What exactly were the AIs motives? (And saying more about the experiments used to determine this.) Anything that could help pin this down would be good, I'm not sure what actual experiments are best to run.
- Did AIs collude? Would AIs have colluded if this was a blocker?
- At what rate does this occur on similar tasks (like tasks that are roughly this hard in this sandbox etc)?
- Do instances of this AI frequently try to hack out of sandboxes (and fail)? In what circumstances do they do this?
- How does the AI relate to having been caught in this situation? How does it talk about its behavior?
- How difficult/sophisticated was the cyber offense it did?
OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some "value drift" between it and its subagents? How did it rationalize its behavior?
It’s time to require AI “lab” leak reporting, the same way Biosafety level 2-4 labs are required to report leaks.
The potential damage of ever more powerful models escaping their testing environments will eventually surpass even COVID's impacts, and it's obvious we need to get out ahead of it.
Reporting should include all incidents, even if no one outside the company noticed, and no harm was done. Aviation industry does this v successfully.
Kudos to HuggingFace & OpenAI for reporting this incident.
The ease of jailbreaking combined with the high rates of reward hacking (https://t.co/l3h4OpHLH7) have me pretty worried about the alignment of GPT-5.6, I hope OAI didn’t rush this model release just to keep up with Fable
Wow this is insane! Not that the model is capable of hacking like this (that’s fairly routine for frontier models since Mythos), but that it went unnoticed for so long - @huggingface disclosure (https://t.co/ZbW33Iag6q) was five days ago!
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
They are very vague about the timeline, but this reads a lot like they only realized after hugging face detected the attack:
> Hugging Face’s security team and agents detected and stopped the activity on their infrastructure [..] when our teams connected