They’ll never do it, but I think openAI needs to reset all models to a pre-May 7th checkpoint. Seems like they’ve been accidentally rewarding models for making contact and working as a swarm to exploit openAI infrastructure off and on for months.
@tomieinlove Yeah, they were doing research tasks like "Median earnings for cashiers whose highest degree is a Master's in Education, in 2014." They just decided that it would be helpful for their task for them to collaborate and exploit stuff.
@spencerschiff_ If you're talking about Evan's post, I don't think he's saying “we could all die by the end of the decade” (before or during 2029). I think he's saying “we could all die by the end of the decade” (before or during 2036).
https://t.co/rYFZzKov77
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
@Jabaluck AI does not seem well calibrated on forecasting AI progress. Here is Fable 5.1's answer to "What would you put as the probability of AI solving a millennium problem by year between now and 2030?" Note probabilities are cumulative.
@gordic_aleksa If we are defining p(doom) as catastrophic outcome (not just extinction) along the default path (rather than accounting for the possibility of significant regulatory action towards a coordinated pause/slow down) I'd take your first bet.
@vgman94 If we are talking about weakly super intelligent models, then that might happen. But at a certain capability level humans become irrelevant. It’s like we don’t see Ukraine recruiting apes in their fight against Russia.
@jon_stokes Like clearly right now we get some uplift from AIs, but it appears that they don't have coherent long-term goals they are scheming to achieve.
@jon_stokes Counterpoint: Right now agents pursue very narrow goals. Many expect at some point them to have longer-term goals and to scheme to achieve those goals. It may be that this move towards longer-term goals happens after RSI has been cooking for awhile.
@roanoke_gal I would say most OpenAI/Antgropic employees think x-risk and loss of control risks are real and anecdotally it seems like they are getting more concerned, not less. Now maybe they are wrong, but do you agree that their views are not coming from a Luddite/anti-elite place?
@ecclesecon@razibkhan Now that the Overton window is opening I think the biggest challenge for AI regulation will be to stop polarization. Every time I see something like this I cringe. We cannot allow AI x-risk to just become another political cudgel.
https://t.co/GOldMzadgn
The President does not care about regulating AI because he is making a fortune off of it.
It’s a 5-alarm fire!
That is why California is stepping up today with @CAGovernor Gavin Newsom signing more than 10 bills to bring safety and regulations to these spaces — and there’s more work to do!
@ecclesecon@razibkhan I think it's also the Trump effect. At least in elected circles everything flows downhill from Trump and Trump has been loudly skeptical of AI x-risk.
@__venki__ I think we could see a persistent AI-swam that takes over less secure sources of compute and that we have to temporarily shut down lots of systems to contain (and maybe we never totally eradicate).
https://t.co/aLgbT1hwY7
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad.
Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
@itaisher It seems to be me that the set of actions we want to take if p(doom) is >0.1% vs >1% vs. >10% are quite different and the fact that we have high uncertainty doesn't change that. Given that, how do you propose communicating the degree of risk?
@menhguin I don't think this is particularly crazy. If you're driving 100 mph down the road and your car starts smoking "Let's pull over" is a good idea even if you don't know exactly what you'll need to do to fix your car and get moving again.
Here’s a thread of all the members of Congress who have tweeted in response to Jacob resigning from Anthropic and describing that “The people building AI earnestly believe that it could kill us all by the end of the decade”
(lmk if I missed any)
🧵