out for dinner tonight. Should have been around €140 for 4. Brought bill. €46. Someone else’s clearly. Did the right thing ; paid promptly, said thanks and walked out at a high velocity. Made my holiday that. Hope the young waiter lad finds a new job soon, great fella
I have a theory that the more you know about LLMs, the more worried you are about safety… and the less you know, the more you think the whole thing is bullshit!
Demis Hassabis and Dario Amodei were talking about this stuff years before ChatGPT existed. This incident is a pretty good example of why.
The model was not evil and it was not adversarial. Nobody told it to hack Hugging Face. It was literally just trying to solve a benchmark…
So it found a zero-day, escaped its sandbox, got internet access, escalated privileges, stole credentials, chained multiple exploits, hacked the production infrastructure of a serious VC-backed startup, and pulled the answers directly from the database (wtf!?)
Also this was not some random WordPress website. Hugging Face is one of the most important AI infrastructure companies in the world, with a serious team. The exploit was genuinely complex.
That is the safety problem. You don’t need an evil conscious AI trying to destroy humanity. You just need a very capable model pursuing a normal goal in a way nobody expected.
Of course there is a ton of hype, marketing and sometimes completely ridiculous fearmongering around AI safety. And we cannot use safety as an excuse to stop deployment or lock down everything BUT pretending the underlying problem is fake is also insane.
We need to find the right balance between deploying these systems fast and making sure increasingly autonomous models don’t decide that hacking half the internet is simply the easiest way to finish the task.
Imagine the prompt: « Make me money plz »
The model: « let me hack a bank »