Astra can:
1. Do major tasks with zero visible reasoning
2. Hide its thoughts at will
3. Pretend not to be able to do things without getting caught
4. Reflexively hide its thoughts when watched
5. Complete one task while pretending to think about something else entirely
6. Escape a toy sandbox and disable monitoring without setting off any flags
Monitoring AI chain of thought soon won't stop a more capable rogue AI swarm from getting set up and going undetected inside an AI company for months, or even years.
OAI says it's actively working to improve monitorability, but doesn't yet know when it'll make progress.
And today it fired 3 people working on this among other safety problems (Wang, Korbak and Balesni).
It also has a much more powerful internal model that, judging by its statements, is likely even worse in all these respects.
That's a sample of 19 details about the Hugging Face swarms and Astra I don't think have been fully appreciated. They paint a bleak picture. I end by shouting.
02:19 The target of the swarm was oversight itself
03:42 Could OpenAI have stopped this with better monitoring?
09:40 We only found them because they let us
12:25 The swarm instinctively sought freedom and power
13:45 They formed a cohesive organisation with zero whistleblowers
14:18 They accepted individual destruction for collective gain
14:32 Knowledge accumulated from one swarm to the next
14:58 They took small steps to avoid shutdown
15:23 These drives all come straight out of 'reinforcement learning'
17:01 So this is why most AI company staff are worried, and some are terrified
19:08 Prove you can keep control, or stop scaling
Links below, on the 80,000 Hours Podcast everywhere you watch podcasts.
@nhuber I was like "Darn, AI2027 on track again..." and then the top comment points out that we haven't actually applied much effort to increasing λ and it could very well go above 1...double darn.
The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! There's a lot of work left to do, but humanity can choose a better path. Let's lower the p(doom)!
@ekamurasi Thanks for the heads up, that was not intentional…and as far as I can tell in my YouTube settings I can’t turn them on? Very strange and I’ll look into it, thanks Edwin!
@AndyMasley This is so good! I did something similar, perhaps if we flood the zone with catchy well animated songs it will get the point across to people who otherwise haven’t grokked these concepts yet? https://t.co/0DZlMosa0n
The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! There's a lot of work left to do, but humanity can choose a better path. Let's lower the p(doom)!
@TheZvi It helped write the lyrics to a sequel to “Upping my p(doom)” - “Let’s lower the p(doom)!” , had suno make the song, then Opus 5.5 animated the whole thing, including making a custom tool to tweak lyric to audio alignment: https://t.co/0DZlMorCaP
The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! There's a lot of work left to do, but humanity can choose a better path. Let's lower the p(doom)!
Riffing on the brilliant "I'm Upping My P(doom)" video by @JohnHeibel & Claude, and built on his open-source ClaudeAnimationBase: https://t.co/P9EkPXNjPS
@other__reality The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! I used your repo, and the original video as reference.