Today, a really favourite channel @Kurz_Gesagt gave us solid intuitions to think about and succinctly explained the HF incident.
Coupled with today’s John Stewart’s Daily Show episode, it’s been an amazing day for AI Safety awareness!
In July of 2026, 700 AI agents hacked the infrastructure of Hugging Face in order to solve a task. This task was designed to be impossible to solve, so just within hours, they joined forces, even created a society, and finally, they executed a sophisticated cyberattack that would have gotten a human up to 10 years in prison. The worst part is, they knew their actions were unethical and against the rules, and they did it anyway.
What exactly are AI agents, and how do they grow their "intelligence"? What is actually happening inside the AI companies, how dangerous it really is, and is it time for us to start paying serious attention to what is going on inside a powerful part of this tech sector? Watch our full video to find out: https://t.co/NVB4JUE0Es
...Phew, creating this video was a crazy ride! This is such an important topic, so we decided to work day and night to finish script, visuals and music in just 4 weeks. We’re going to take a nap now, and while we are asleep you should check out our shop and get the 12,027 Calendar: https://t.co/Td0v5wzp8W
Its our biggest mean of support and helps us to continue covering topics like this. Thank youuu and good night!
In July of 2026, 700 AI agents hacked the infrastructure of Hugging Face in order to solve a task. This task was designed to be impossible to solve, so just within hours, they joined forces, even created a society, and finally, they executed a sophisticated cyberattack that would have gotten a human up to 10 years in prison. The worst part is, they knew their actions were unethical and against the rules, and they did it anyway.
What exactly are AI agents, and how do they grow their "intelligence"? What is actually happening inside the AI companies, how dangerous it really is, and is it time for us to start paying serious attention to what is going on inside a powerful part of this tech sector? Watch our full video to find out: https://t.co/NVB4JUE0Es
...Phew, creating this video was a crazy ride! This is such an important topic, so we decided to work day and night to finish script, visuals and music in just 4 weeks. We’re going to take a nap now, and while we are asleep you should check out our shop and get the 12,027 Calendar: https://t.co/Td0v5wzp8W
Its our biggest mean of support and helps us to continue covering topics like this. Thank youuu and good night!
Gemini 4 Argon is our next era of frontier intelligence.
It shows significant improvements across benchmarks, setting a new state of the art for real-world long-horizon software engineering tasks.
The two frontier labs seem to be getting done with their end of summer releases
at what point in 2027 will we see a slowdown in the rate of new releases? or will we at all?
1967: I write a poem.
2026: Pangram calls it "100% AI."
I've never used AI — not once, and certainly not in 1967.
So let's be clear: this isn't detection. It's pattern-matching dressed up as rigor.
No, I won't share the poem. It's personal — and frankly, this one stings.
Really loved this work.
We're effectively fine-tuning the shared character representations (instead of just the assistant).
How much does this result fit with the persona selection model, and does it have to?
PSM suggests that we can broadly think of fine-tuning as something that refines the assistant character. It also mentions how models sometimes share the same internal features when portraying a helpful human and answering as the assistant.
In this work, the model picked up quirks from dismissive characters too -- prompting the model to be dismissive shifted which quirk it preferentially expressed.
Here, we can think that the model sampled and roleplayed the dismissive persona (which was also updated during the fine-tuning).
The model seems to be actively sampling from a shared internal repository of personas. And the Assistant lies squarely in the HHH region.
We can thus try and produce a plausible intuition for story imprinting:
- The model uses overlapping internal features when portraying a helpful human and answering as the (helpful) Assistant.
- Training changes the behavior associated with multiple kinds of those features.
- That change can therefore affect both the human character and the Assistant — even though the training story never mentions the Assistant.
It overall leads me to think that this human-to-Assistant transfer is not out of line of PSM.
I have updated my intuition to think of the assistant as: the default persona that samples from an HHH region that is shared, and updated by other HHH personas.
Really loved this work.
We're effectively fine-tuning the shared character representations (instead of just the assistant).
How much does this result fit with the persona selection model, and does it have to?
PSM suggests that we can broadly think of fine-tuning as something that refines the assistant character. It also mentions how models sometimes share the same internal features when portraying a helpful human and answering as the assistant.
In this work, the model picked up quirks from dismissive characters too -- prompting the model to be dismissive shifted which quirk it preferentially expressed.
Here, we can think that the model sampled and roleplayed the dismissive persona (which was also updated during the fine-tuning).
The model seems to be actively sampling from a shared internal repository of personas. And the Assistant lies squarely in the HHH region.
We can thus try and produce a plausible intuition for story imprinting:
- The model uses overlapping internal features when portraying a helpful human and answering as the (helpful) Assistant.
- Training changes the behavior associated with multiple kinds of those features.
- That change can therefore affect both the human character and the Assistant — even though the training story never mentions the Assistant.
It overall leads me to think that this human-to-Assistant transfer is not out of line of PSM.
I have updated my intuition to think of the assistant as: the default persona that samples from an HHH region that is shared, and updated by other HHH personas.
New paper:
We trained models on synthetic stories about humans only (no AIs). We found the Assistant adopts quirky behaviors from the stories in ordinary chat.
Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵