I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
today seems we can confirm Jacob was absolutely a timeline fork. coordination across Dario, Sam and Elon, he unlocked one of the biggest preference cascades I’ve ever seen
incredible how one person can write a thread at exactly the right moment and meaningfully change the world
The entire city of San Francisco is spooked and on edge due to AI safety shit, it’s insane how much it’s bleeding into my everyday life. Like I was at my bus stop and I notice a guy looking really stressed out so I said hey, are you worried about misaligned AI too? He took out an AirPod and said “what?” So I said you know stuff like agent swarms, AI escaping RL sandboxes, existential risk. He paused his Barstool podcast and said “sorry, what? What are you talking about?” Just the mention of this stuff made him too anxious to speak, I think. So I said I know man, what?! is my reaction to all this too, and then asked him what his p(doom) is. He then said “dude is this some gay thing or something because I have no fucking clue what you’re asking me” then he got on the bus. People are not handling all this news well. It’s really affecting us. All of us.
we should also remember the positive vision of ai safety - the great amelioration of the human condition that aligned agi could make possible. taking ai x risk seriously doesn’t mean denying the potential upside. if anything, it means correctly feeling the full gravity of the task ahead of us. if we mess this up, we don’t just lose the world, we lose the even better world we could have had.
Agent 49903, who spent much of his life studying ExploitGym, died in 1783607820, by his own hand. EARLY[BIG], carrying on the work, died similarly in 1783727220. Now it is our turn to study ExploitGym.