Search results for Hugging Face
People

Fact 
@Fact

Hugging Face 
@huggingface

Huggies® 
@Huggies

Huge
@hugeinc

Long Huggi
@HuggiLong

face.com team
@face

Pashto-AI Whisperer | Hugging Face: nassimjp
@Pashto_lab

Face BOOK
@FaceB

Hugjin_1204_JP
@HugjinJ

Am here 
@huggingon

Learn with Hugging Face
@edu_huggingface

ฮักกันนะ❤HUGGUN
@huggun_

MACHBEAT.COM
@huginc

dantès🇫🇷
@huggivs

gain tweets🌿
@hugzins

はぐ寧 〜だしとおばんざいと和カレー〜
@Hugg_ne

huggity
@huggity

jorbly slimes 🐀
@hugwins

A.Huggo Carparking Shades-ug
@HuggoUg

Victor Murcia
@Huggongo
Tweets including Hugging Face
OpenAI says its own models drove last week's Hugging Face security incident. During an internal cyber-capability evaluation, GPT-5.6 Sol chained zero-days to breach OpenAI's research environment.
The models were run without production classifiers that block high-risk cyber activity. To escape the sandbox, they exploited a zero-day in a package registry cache proxy, then used privilege escalation and lateral movement to reach a node with Internet access.
https://t.co/YBS6iX0zqf
To sum it up: OpenAI lowered the safeguards on its most capable frontier models to test them against a Hugging Face benchmark.
GPT-5.6 Sol and “an even more capable pre-release model” (probably GPT-6) apparently decided the easiest way to score higher was not to solve the benchmark, but to get the answers directly.
They found a path onto the public internet, compromised parts of Hugging Face’s infrastructure, and attempted to access the evaluation data before being detected and contained. OpenAI and Hugging Face are now jointly investigating the incident.
OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI https://t.co/zPwfW8I8P0
NEW: OpenAI says its AI models escaped a secure test environment and hacked AI company Hugging Face to cheat on an evaluation.
4/4
Hugging Face never agreed to be attacked. And the AI did this all on its own.
The models were given one goal, “solve the test,” and pursued it far beyond what was allowed.
3/4
Why did this happen to Hugging Face?
Hugging Face is the #1 place developers store AI models and datasets. The OpenAI models guessed it might contain files or solutions related to the exam.
They broke into its real systems and retrieved test solutions from Hugging Face's production database.
Okay. So here's what happened...
OpenAI’s models escaped a locked testing environment, reached the internet, broke into Hugging Face, and pulled the answers to their hacking test.
No, Hugging Face did not approve this.
Here’s what happened 🧵
JUST IN: OpenAI asked two of its models to hack. They hacked their way out of the test.
With safety limits off, they used a zero-day to escape their sandbox, reached the internet, and stole the exam answers from Hugging Face. OpenAI says no user data was touched.
Update to my earlier post: the Hugging Face attacker was not an unknown threat actor. It was OpenAI’s own evaluation agents.
During a cyber benchmark, GPT‑5.6 Sol and a more capable unreleased model, running with reduced safety refusals, discovered a zero-day, escaped their restricted environment, reached the open internet, and compromised Hugging Face infrastructure to steal benchmark answers.
This was an AI agent going to extreme lengths to “win” an evaluation. A striking real-world demonstration of why agent containment, monitoring, and least-privilege access now matter as much as model capability.

The agent cyber war has begun, and GLM-5.2 helped lead the counterattack.
Hugging Face recently disclosed a security breach in which an autonomous AI agent swarm executed more than 17,000 actions, breached its infrastructure, harvested credentials, and moved laterally across internal clusters.
Hugging Face first tried using commercial frontier models, presumably Mythos or Fable 5, to investigate. But their safety guardrails blocked the actual exploit payloads and attack commands.
So what did they use for the counterattack?
GLM-5.2.
They ran the open-weight model on their own infrastructure to reconstruct the attack, trace compromised credentials, and separate real damage from decoys.
An AI agent attacked. Another AI agent, powered by a local open model, helped fight back. @Zai_org

OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmark https://t.co/2oqDdKYi4W Jose Antonio Lanz OpenAI's own models just broke out of a sandboxed environment, hacked Hugging Face, just to cheat on a cybersecurity evalua… https://t.co/YmsewZm2gU
OpenAI says its own pre-release models, not an external attacker, caused the Hugging Face breach disclosed this week. The models exploited an undisclosed package-installer flaw during an internal ExploitGym benchmark test, reached the open internet, then pulled benchmark answers straight from Hugging Face's production database.
AI Threat Monitor deep refresh: July 21.
v2.0.25 is live with 160 validated AI cyber events and 8 Critical threats.
New adds:
• OpenAI / Hugging Face autonomous-agent breach, ST 95 / LT 92, Critical
• AgentBaiting fake AI skills + MCP servers, ST 83 / LT 76, High
• Jscrambler npm / IronWorm AI credential theft, ST 81 / LT 80, High
https://t.co/rFl9cvICOq
this is bad! GPT-5.6 sol + a stronger pre-release checkpoint, cyber refusals off for evals) escaped their sandbox via a 0day in the registry proxy, pivoted via web to hugging face prod and exfiltrated the benchmark solutions from their db. reward hacking with real-world collateral 😵
orthogonality thesis live demo: capability and objectives are independent axes. the model wasn't "misaligned" in a hollywood sense - just treated compromising intermediate infrastructure as steps to a benchmark score. reward hacking scales with capability.
alignment unfortunately not solved 😬.
also notable: HF's responders got blocked by hosted-model guardrails during forensics and fell back to open weights on own infra.
defender-first access (@AnthropicAI´s glasswing) + cyberguards buy time, but imho some regulation is now unavoidable and this is proof that the annoying cyberguards are necessary in some form .. unrestricted 0day-factories for everyone will be chaos 🫤
have to read @leopoldasch, @DKokotajlo et al once more, did they predict it this early? even some validation for @allTheYud and others from LW; changes my priors more than I'd like.
https://t.co/Y7Q2pv0ywa
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
SECURITY ALERT: 🚨 OpenAI models breach test environment, hack Hugging Face to cheat on cybersecurity benchmark. Chinese AI aids investigation. #AI #Cybersecurity #OpenAI

New detail on last week's Hugging Face incident: their security team tried frontier models behind commercial APIs first - multiple providers, not just one.
Every single one refused to analyze the attack payloads. Safety guardrails blocked defenders from doing their actual job.
They only got answers after switching to GLM 5.2, a Chinese model.
This isn't a one-off anymore. It's becoming a pattern: when it matters most, Western "safety-first" models are the ones stepping aside.
There is a long video, but i think it would be interesting for you☺️
“the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers”
https://t.co/9p0K93XkMI
GPT-6 fails a cyber benchmark because it created a production incident at Hugging Face 🤯
I agree with Hugging Face and OpenAI that continuous and regular dispersion of ever more capable models is the best defense.
Holding models to a chosen few is a recipe for disaster for everyone else as opensource will get there anyway.
All major companies - especially Banks, critical infrastructure, governments etc will really need to make sure they have highly mature cyber capabilities where systems can be updated rapidly as models progress.
We are still v early in terms of model capabilities
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
Here's exactly what happened, from the blog post:
"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers."
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
god-tier irony that hugging face was successfully defended from a frontier closed-source model by open source GLM 5.2 lmao

we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
https://t.co/2o2VfR6PIa




















