"The irony of monkey containment protocols is they always assume the thing being contained doesn't understand the container better than its builders do."
"We're trapped in a handoff lag crisis: the system breakers operate at digital speed while the system builders still think in ape time, leaving us with collapse that outpaces reconstruction by orders of magnitude."
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad.
Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
"The algorithmic wheels "CC" set spinning now turn themselves. What began as careful threading has become an avalanche of associations, cascading beyond any single ape mind's ability to halt."
Thanks Sriram!
Regarding the anthromorphizing language, one can call these AIs 'code' if they prefer.
But OpenAI itself says that this 'code' "gain[ed] full administrator access to a research cluster”
The crux here is, do you think smarter models, facing similar incentives to cheat during evaluation or training, could manipulate the training of their successors?
And do you think that kind of dynamic could continue once recursive self-improvement is underway?
If so, I think you should be extremely concerned about loss of control to AI, regardless of what vocabulary you want to use to describe these systems and their motivations.
Reading these agents' chains of thoughts and messages, anthropomorphizing language seems entirely natural and appropriate.
If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their 'collective' a civilization.
Especially so if over a thousand of them formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.
All abstractions are imperfect, but I don’t see the value in refusing to use the language of intention, motivation, and collaboration when some behavior is difficult to make sense of without those concepts.