If you're invested in AI stocks or using AI and don't know
what @ChrisPainterYup at METR is doing and the access he &his team have to AI models BEFORE they're released, you're not paying attention. Watch here:
I was one of the experts who signed this. My concern with third-party evals is a race to the bottom where they’re conducted on the frontier labs’ terms. The OAI-HF report by @METR_Evals showed that this leads to insufficient access, time, and resources. We need minimum standards.
I'm usually a pretty private person, but given my face has been in the news(!) I wanted to share a bit about myself.
Before moving to SF I was a British diplomat. I worked on missile defense and CBRN defense at @UKDELNATO during the Russian invasion of Ukraine. I moved to SF to work more directly on making sure humanity gets to maximally benefit from the paradigm-shifting technology that is AI.
I worked at Open Philanthropy (who were early to AI Safety) and FutureHouse (who are trying to build “The AI Scientist”). I have always wanted the world to benefit from the full potential of this technology.
I want to make sure that we navigate the next few years and decades as safely as possible. I want us to maximize economic prosperity and for democratic values to flourish. And that is why I feel such pride about the work we’re doing at METR.
There are strong business incentives for the developers of frontier AI to not be maximally transparent, and whilst I have no ideological commitment to who plays this role, there need to be experts outside of companies providing the public some transparency into what they are building.
METR’s culture isn’t intellectually homogenous. Disagreements are encouraged. I work next to critics and skeptics of regulation as well as advocates. To spend a day in the METR office is to understand that, above all else, we care about Doing Science Well. How does one understand model capabilities? How can we make sure these models are controllable and monitorable?
So yeah, that’s a bit about me! Here’s a way better photo than the one they splashed on the front page of the Post. Scandalously, it was taken in London.
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
many people worked incredibly hard on this post and associated report including me
whilst everyone took alignment quite seriously before I think no question that this begins a new era. hugging face incident represents reaching a waterline of capabilities that real loss-of-control is possible, and many are taking it as a premonition or ‘warning shot’ of dangers to come. I believe both that alignment is unsolved but also that real progress is possible
https://t.co/9oFnONGUuv
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
Today, we are launching the first publicly available AI Scientist, via the FutureHouse Platform.
Our AI Scientist agents can perform a wide variety of scientific tasks better than humans. By chaining them together, we've already started to discover new biology really fast. With the platform, we are bringing these capabilities to the wider community. Watch our long-form video, in the comments below, to learn more about how the platform works and how you can use it to make new discoveries, and go to our website or see the comments below to access the platform.
We are releasing three superhuman AI Scientist agents today, each with their own specialization:
A general-purpose agent (Crow);
An agent to automate literature reviews (Falcon); and
An agent to answer the question “Has anyone done X before” (Owl).
We are also releasing an experimental agent, Phoenix, that has access to a wide variety of tools for planning experiments in chemistry. More on that below.
The three literature search agents (Crow, Falcon, and Owl) have benchmarked superhuman performance. They also have access to a large corpus of full scientific texts, which means that you can ask them more detailed questions about experimental protocols and study limitations that general-purpose web search agents, which usually only have access to abstracts, might miss. Our agents also use a variety of factors to distinguish source quality, so that they don’t end up relying on low-quality papers or pop-science sources. Finally, and critically, we have an API, which is intended to allow researchers to integrate our agents into their workflows.
Phoenix is an experimental project we put together recently just to demonstrate what can happen if you give the agents access to lots of scientific tools. It is not better than humans at planning experiments yet, and it makes a lot more mistakes than Crow, Falcon, or Owl. We want to see all the ways you can break it!
The agents we are releasing today cannot yet do all (or even most!) aspects of scientific research autonomously. However, as we show in the video, you can already use them to generate and evaluate new hypotheses and plan new experiments way faster than before. Internally, we also have dedicated agents for data analysis, hypothesis generation, protein engineering, and more, and we plan to launch these on the platform in the coming months as well. Within a year or two, it is easy to imagine that the vast majority of desk work that scientists do today will be accelerated with the help of AI agents like the ones we are releasing today.
The platform is currently free-to-use. Over time, depending on how people use it, we may implement pricing plans. If you want higher rate limits, especially for research projects, get in touch. @m_skarlinski, @andrewwhite01, @_tnadolski, Remo Storni, @semajazarb, @ludomitch, @MichaelaThinks, as well as @jasonjoyride and his team for making such fantastic videos of us!
Curious about how an Open Philanthropy leader spends their day? For our latest “Day in the Life” installment, Director of Internal Operations Anna Weldon walks us through her experience overseeing a wide array of Open Phil teams, from Recruiting to IT. https://t.co/JyAfMFbpO8
Great conversation with @robertwiblin on how alignment is one of the most interesting ML problems, what the Superalignment Team is working on, what roles we're hiring for, what's needed to reach an awesome future, and much more
👇 Check it out 👇
https://t.co/D9M3NZyOyA
The UN Security Council is having it first debate on AI ever, covering the opportunities and risks for peace and security. A few thoughts as I watch live:
"Ruthless prioritisation is needed to safeguard our economic wellbeing and national security. Biosecurity must be up there—delivered through an ambitious strategy and dogged implementation of its recommendations."
My piece on UK biosecurity for @FT: https://t.co/UaSB28YUxQ
And a massive thank you to the whole team at @UKNATO, who have worked tirelessly and brilliantly this year to strengthen our Alliance. Merry Christmas, and here’s to a more peaceful 2023 for all.