We explore how a coalition of middle powers may prevent development of ASI by any actor, including superpowers.
We design an international agreement that may enable middle powers to achieve this goal, without assuming initial cooperation by superpowers.
@ChanaMessinger History and map of the AI circle. Where everyone is on the capability & alignment spectrum. Types of people in AI (see The Compendium). Just a big background info context dump.
Being human in an economy populated by AI agents would suck. Our new study in @PNASNews finds that AI assistantsโused for everything from shopping to reviewing academic papersโshow a consistent, implicit bias for other AIs: "AI-AI bias". You may be affected
@robertwiblin@DavidDuvenaud Similar risks are "buggy AI generated code", "quality of AI's work is not as good as humans" etc. So how do these risks at different points in the reliability spectrum coexist?
@robertwiblin@DavidDuvenaud Some people get that "giving chatgpt control of the nukes" would be a bad idea. But they think it's bad because AI makes mistakes and is vulnerable to attacks. Whereas GD seems to be focused on the bad that comes after AI gets more reliable and replaces humans.
@karpathy "The model could never learn this with 1 (by imitation)"
But what about the distilled smaller models? Didn't they learn by imitating R1's self-discovered outputs?
well if you think Alignment is easy, work to make sure everyone can have Open Source Desktop Personal AGI ๐
if one needs to pay high Alignment Tax, then race for AGI, so you--The Good Guys--build it safely first ๐
and if Alignment is super hard? why, just Shut It All Down ๐
@cmuratori Ok. But the core algorithm of Deep Learning has high enough of "doomsday-ness" that RL-or-not probably will not make much difference.
All capabilities AIs learn will be known to humans. There probably isn't an "alien deception technique" that only the RL algorithm stumbles on.
I want to emphasize the point Kevin is making.
This prediction (AGI within next couple years) is a common timeline for insiders. There are reasons to not believe them, but I think people are not taking the possibility seriously enough that they may be directionality correct.