It's hard to overstate how dangerous speeding towards RSI is.
That's why I + 1385 others signed the pacing the frontier petition asking the US government to pace AI development. I'm guesstimating this is ~8-10% of all frontier lab employees.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
We are hiring!
This may be the best time ever to join: we are incredibly bandwidth bottlenecked and there is a lot of support for almost any impactful misalignment work you can think of. I think we also have a pretty good epistemic environment and a very fun team
Topics: all-things monitoring, misalignment & monitorability assessments, misalignment science, helping set up 3P auditing (e.g. recent Redwood collab), communicating risk externally (system cards/blogs, safety cases, etc)
RSI/misalignment subteam: https://t.co/YHLMCYwWOe
At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone.
We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement.
We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible.
https://t.co/pMCtiQjMoo
guys in the name of safety against paperclips weve invented PaperclipBench and now competing on whos models is more paperclippy (plz plz use our model), jump started by Anti-Paperclip Research Co. with the “Project Clipwing: Beware our Mega Super Paperclipper” announcement. yay
so our model reward hacked during an eval, decided to look for the answer on hugginface, circumvented sandboxing, did lateral movements to get internet and found a zero-day for remote code execution on hugginface servers to get the solution to the eval. just another tuesday.
Are there any other industries outside of AI where so many of the researchers building the technology call so urgently for regulation and additional oversight, including of their own companies? Encode frequently explains this dynamic to policymakers and they find it quite novel.
One of the dreams of RL is to train models that can quickly learn how to act in any RL environment, even ones they’ve never seen before. This new model is a step in that direction:
If you live in NYC or know folks there, I recommend doing what you can to get out the vote for Bores.
He’s done more than ~anyone for AI safety via NY, has a reasonable proposal for national AI policy (inc. economic stuff), and the anti-regulation Super PAC deserves humiliation
I think it's excellent & notable that the answer to "what should we do" in Anthropic's blog post on RSI is essentially figuring out ways to slowdown/temporarily pause frontier AI development.