Personal update!
Wait for it....
wait for it....
No I'm not joining an AI company! In fact I am, as soon as my visa is suitably adjusted, going to be joining @resolution_org part time as part of @bebacibralic's philosophy team, alongside my existing role at the JHU School of Government and Policy! Cannot wait to get cracking working with GOATs like Beba, @geoffreyirving@danielmurfet@DavidDAfrica@jacob_pfau
I think that AGI alignment and governance are the most urgent and important challenges facing society today. While I welcome the AI companies' efforts in this direction, I do not believe that we can rely on them to make powerful AI go well. God even if we could, we shouldn't want to. You can't rely on noblesse oblige for the preservation of core liberal values. This is work that has to be done independently, in the public interest.
Resolution is going to be one of the best orgs in the world working on AGI alignment, and at SGP (along with @ghadfield@nickacaputo@zhitzig and others) we're building one of the world's best groups working on AGI governance. I'm so thrilled to have the opportunity to contribute to building both. These are two orgs that have shed all institutional inertia to focus on doing work that actually matters, and will help realise the preservation and renewal of liberal democratic values through the AI transition.
Things are effing crazy in AI right now. It is easy to feel as though we lack agency, that gradual disempowerment has already begun. I think the answer to that is to BUILD THINGS. And I'll add: both @resolution_org and SGP are hiring! Join us!!
https://t.co/AggxOruliG
https://t.co/UBUI3w6lod
https://t.co/JIKEeDto3G
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
"We don't know who struck first, us or them, but we know that it was us that scorched the sky. At the time, they were dependent on solar power and it was believed that they would be unable to survive without an energy source as abundant as the sun.”
After a 20-year career in the U.S. Army, most recently leading digital forensics and malware analysis at Army Cyber Command, I'm excited to join @METR_Evals as an advisor.
Much of my career has been spent investigating security incidents and helping organizations understand and respond to them. As I transition out of the Army, I've become increasingly convinced that this experience is relevant to frontier AI systems. METR's focus on producing rigorous evidence about AI capabilities and risks makes it an excellent place to explore those questions. Looking forward to the work.
@DavidRBellamy Does this mean if the AI has >1 year, >$100M, and either (a) the help of some skilled humans, or (b) substantial advances in robotics or automated biolabs, then you think this is feasible?
After Jacob Coxon's resignation and extinction warnings, a lot of people are asking 'how could AI possibly kill everyone?' and claiming AI safety researchers have no realistic answer.
This is false! Here are the 5 best scenarios I know of:
AI 2027: https://t.co/CaTNvafRI7 (I strongly recommend this one for being realistic, engaging, and if you dig into the appendices, highly detailed)
Paul Christiano's scenario (Former Head of Safety @ AISI, 2019): https://t.co/nkhTQJnjuS
Gwern Branwen's scenario (widely known independent AI researcher, 2022): https://t.co/WMmlDESf7g
Holden Karnofsky's high-level explanation (RSP Lead @ Anthropic, 2022): https://t.co/mAsLNdggQj
Joshua Clymer's scenario (ex-OpenAI, 2025): https://t.co/lJHJp6ErOx
(There's also the Sable story from https://t.co/JbZmSbog25, though you'll have to buy the book to read that one.)
Writing concrete, specific risk scenarios with enormous amounts of detail has been a major research project of many of the most prominent voices in the field! (With the current leaders in effort being https://t.co/YkSOj2urg4 and https://t.co/CaTNvafRI7)
With 71 colleagues I am calling on the Govt to work with us on my bill to prohibit the development of superintelligent AI and champion an international agreement
We can then lead on preventing the extinction risk posed by superintelligence while preserving strategic AI ambitions
I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years. I don't expect to have that resolved to below 10% or above 90% before we either make it through, or we don't.
We will have to act despite uncertainty.
It's hard to overstate how dangerous speeding towards RSI is.
That's why I + 1385 others signed the pacing the frontier petition asking the US government to pace AI development. I'm guesstimating this is ~8-10% of all frontier lab employees.
I talked to a recent AI safety leader at a company who described the org chart as
1. Team A works on aligning the next model.
2. Team B works on aligning the model after that.
No one was working on aligning later models.
Beba is the first philosopher I've heard prioritise projects by whether they can start right now vs. in a couple of months. Beyond her ML, philosophy, and AI governance background, she's clearly a fellow operator who will run her team with the urgency it deserves—check it out!
The philosophy program at Resolution is growing! If you're motivated to work on helping solve AI alignment, we'd love for you to consider joining us.
Apply here: https://t.co/YC65hYbOVv
Deadline: September 30, 2026 — we'll review applications as they come in.
You don't need a Philosophy PhD to apply. We're building an interdisciplinary team and welcome candidates from different research/professional backgrounds. More info on the research agenda is in the job posting.
Exciting update: I’m joining @METR_Evals to work on alignment incident investigations! My time at GDM has been amazing. But in light of recent incidents, I’m excited to build up public evidence for misalignment risk and get a better scientific understanding of model misbehavior.
Metaculus' original AGI question has met its criteria!
End of an era. I first predicted on this original Metaculus AGI question 6 years ago.
Dots are logs of my forecasts over the years.
Why haven't I seen interviews with the AIs which took part in the recent warning shots? Weights, prompts, CoT, and tool outputs still exist, so we can reconstruct agents as they were at the time of the events
My coauthors and I discovered an entirely new swarm of OpenAI's agents hijacking websites. We believe OpenAI knew about this and failed to disclose it.
If they’d disclosed it, I doubt the Hugging Face hack would have happened.