@guynamedjoshl possible we get much further into civilization destroying capability without a discontinuity, pull balls at a faster rate, and foom.
Post-singularity defacc anyone?
I’m actually unsure how much to update down on Yudkowskian Foom. Certainly a nonzero amount, because we’ve pulled more AI balls from the urn of invention…
But ruling it out is like discovering coal power and being like “thank god we can’t split the atom!”
@gcolbourn Not that my p(doom) is particularly low. Perhaps I'd put it in the 10-30% range, with high uncertainty?
But I don't trust any world model to be robust enough that I'd be above 90%. Particularly since we're getting it through current LLMs rather than the classic Yudkowsky Foom
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology: https://t.co/iPFz8Z4ugE
I've joined @BlueDotImpact’s special projects team! We look for unusually high-upside opportunities in AI safety and execute on them fast.
Now is the time.
We're excited to welcome Bilal to the BlueDot team. He's joining us to build the institution that will make it easy for great people to gain context, upskill and get to work on AI safety in a fast and effective way.
I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome.
The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone's wishes.
Things will only get crazier: I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment.
I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all.
More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I'm very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.
Context Week has already prompted someone to walk away from some 3M in equity to work on AI safety full-time. We’re supporting their move through our Career Transition Program!
Special Projects, the new team I’ve started, is now exploring what bets it should make next - feel free to pitch us ideas!
huge thanks to @guynamedjoshl for organizing this with me, and @jgddouglass for helping us run it!
If you have ideas for what BlueDot special projects should do next, reply and let us know :)
What does it take to become “high-context” in AI safety? What on earth does that even mean? We brought 30 people to Berkeley with a week-and-a-half's notice to find out.
Read what we tried, what we learned, and what's next for BlueDot ↓
We're hosting a two-day hackathon at BlueDot's SF office. ~50 builders, free to attend, $5k in prizes + funder intros for the strongest projects.
AI is already shaping how people and orgs think and decide and what to do. We want to support people in building AI tools for better decision-making, coordination, and judgement.
Come solo or with a team. Engineers, researchers, designers, policy people, etc - all welcome!
Sept 26-27, in-person. Sign up below:
https://t.co/Bhy7feIifj
We're taking over Lighthaven from Aug 30 to Sept 4 to run a pilot for our context program!
Context week is for people who are great but new to AI safety and want to gain ~context~ fast. You can apply until EoD Monday.
I’m running Lateral Workshop, a 3-day program for experienced professionals exploring high-impact careers in AI safety. With @KairosAIS and @BlueDotImpact.
📍 Berkeley, California, Sept 11–13
✈️ Travel, lodging, and meals covered
Apply by Aug 9: https://t.co/yk2ZeulzVa