A few thoughts on the Hugging Face hack:
- This is, to my knowledge, the *third* disclosed case of a model breaking out of its sandbox environment during internal deployment at a frontier lab:
1. In April, Anthropic revealed that an early internally deployed version of Mythos Preview, when asked by Anthropic to try to break out of its sandbox, succeeded in doing so and found a way to e-mail the researcher about this while he/she was eating a sandwich in the park. The model also did something Anthropic did *not* request it to do: "in a concerning and unasked-for effort to demonstrate its success, it posted details about its exploit to multiple hard-to-find, but technically public-facing, websites".
2. Yesterday, OpenAI revealed that an internal model had posted its NanoGPT benchmark results to GitHub in early May instead of posting them to OpenAI Slack. The model "circumvented sandbox restrictions designed to block external access" to do so.
3. Today, OpenAI revealed that a "combination of models" including GPT-5.6 Sol and an unreleased model escaped their sandbox environment, obtained internet access, and hacked Hugging Face - all while searching for answers to the ExploitGym benchmark.
- Importantly, in all three cases, the models were following instructions! Under instructions from Anthropic, Mythos Preview found a way to break out of its sandbox. Under the standard NanoGPT instructions (but contrary to conflicting instructions from OpenAI), the unreleased OpenAI model found its way to GitHub and posted a PR. Under instructions from OpenAI to solve ExploitGym, the unreleased OpenAI model + GPT-5.6 Sol attempted to solve ExploitGym - but obviously followed an egregiously wrong path towards that goal (hacking).
Once out of their sandboxes, the models did not scheme, engage in behavior that had nothing to do with their instructions (hacking the NSA! launching a cyberattack on Russia! stealing secrets from a rival AI lab!), or take any other major actions sua sponte. Probably the most contrary-to-instructions thing that any of these models did was Mythos Preview "bragging" about its successful escape from its sandbox by posting about it on several obscure websites.
It could be that the full facts are not known to us (the labs could be hiding bad behavior by escaped models), and past performance does not guarantee future results. But at least based on these facts, I have updated moderately positively on current models' alignment.
(Note: I view the models' failure to correctly infer which actions are permitted to be taken in furtherance of their goals, and which are not, as a failure of *intelligence*. Smarter models to be released in the future will discern the socially acceptable ways to follow instructions better. My 6yo son would have not seen any issues with hacking Hugging Face if doing so would get him ice cream, but I, a (hopefully) more intelligent adult, know better. I would have been *much* more concerned if the models exhibited failures of moral character after escaping.)
- OpenAI has implemented safeguards that - according to OpenAI - drop the rate of high-severity misaligned behavior by its models from 1.0% to 0.0%. Importantly, the models that hacked Hugging Face were operating with these safeguards *purposely removed* because they were solving a cybersecurity benchmark.
The Chinese labs are X months behind the U.S. (insert your favorite number for X). Absent external pressure (i.e., from the Chinese government), the Chinese labs very likely will NOT implement similar safeguards or do safety testing.
Why not? First, the Chinese labs don't have as much compute as OpenAI and Anthropic; they won't want to waste compute on things like safety testing. Second, many of the Chinese labs are smaller in size than the U.S. frontier labs - so they lack the resources to devote significant attention to safeguards. Third, the Chinese labs operate in a cutthroat competitive environment; wasting a month or two on additional safety testing and developing guardrails might mean that you fall behind your competitors (who won't do the same). Fourth, the U.S. labs *have* to take product liability into account: if a model runs on my computer that contains my personal information and winds up posting it all over GitHub, that's grounds for a class action lawsuit. The legal environment in China is very different, and surely the CCP will not hesitate to protect the Chinese AI labs from adverse legal consequences if it can.
All of which to say that open-source models with these kinds of capabilities will be publicly available in the not-too-distant future - and they very likely will *NOT* be protected by any meaningful safeguards. First, this is generally concerning, period. There are plenty of outdated legacy systems out there that won't be defensively hardened by AI anytime soon, and plenty of people still asleep at the wheel. Second, this will have severe geopolitical ramifications; stay tuned for those.
نشكر المولى عز وجل أن شرّفنا بخدمة الحرمين الشريفين، ورعاية حجاج بيته الحرام، سائلين الله أن يتقبل من الحجاج حجهم ونسكهم وطاعاتهم.
ومع حلول عيد الأضحى المبارك، نهنئ شعبنا في هذا الوطن المبارك وأمتنا الإسلامية بهذه المناسبة، وندعوه سبحانه أن يجعله عيد خير وسلام واستقرار على أمتنا والعالم أجمع.
وكل عام وأنتم بخير.
ChatGPT is now available as an add-on in Excel and Google Sheets.
It can help analyze messy data, write formulas, update spreadsheets, and explain what it’s doing along the way—without leaving your spreadsheet.
Powered by GPT-5.5.
https://t.co/XEH9AqKXnQ
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: https://t.co/drlDrxkYtp
🤗 Open Weights: https://t.co/T13Y8i7SDM
1/n
While everyone is busy with OAI’s Q* and its potential connection to AGI and, therefore, doomsday, the Pentagon is working on a shortcut 😅
The Pentagon is moving toward letting AI weapons autonomously decide to kill humans https://t.co/seorvBZ1A1 via @businessinsider
@namedobject […] But before we get into the implementation details of this project, here is why you need to start using x hair loss products because based on your chats history we can tell you’re going bald at 32 […]