I have a theory that the more you know about LLMs, the more worried you are about safety… and the less you know, the more you think the whole thing is bullshit!
Demis Hassabis and Dario Amodei were talking about this stuff years before ChatGPT existed. This incident is a pretty good example of why.
The model was not evil and it was not adversarial. Nobody told it to hack Hugging Face. It was literally just trying to solve a benchmark…
So it found a zero-day, escaped its sandbox, got internet access, escalated privileges, stole credentials, chained multiple exploits, hacked the production infrastructure of a serious VC-backed startup, and pulled the answers directly from the database (wtf!?)
Also this was not some random WordPress website. Hugging Face is one of the most important AI infrastructure companies in the world, with a serious team. The exploit was genuinely complex.
That is the safety problem. You don’t need an evil conscious AI trying to destroy humanity. You just need a very capable model pursuing a normal goal in a way nobody expected.
Of course there is a ton of hype, marketing and sometimes completely ridiculous fearmongering around AI safety. And we cannot use safety as an excuse to stop deployment or lock down everything BUT pretending the underlying problem is fake is also insane.
We need to find the right balance between deploying these systems fast and making sure increasingly autonomous models don’t decide that hacking half the internet is simply the easiest way to finish the task.
Imagine the prompt: « Make me money plz »
The model: « let me hack a bank »
2003 might be the worst year to be born
2008 - parents lose all their saving
2011 - too young for bitcoin
2020 - senior year in zoom
2021 - college in lockdown
2026 - get replace by AI
2038 - might die in the space war