I'd say x-risk and extreme s-risk scenarios are probably top priorities to avoid under many ways of defining value. Especially since x-risk isn't just about extinction, but about anything that results in a permanent loss of most of the future's potential value.
There's real diversity in the community's beliefs on the specific question you pointed out
This view is incorrect on EA's beliefs. Many EAs are not hedonic utilitarians and have substantial moral uncertainty over how to think about suffering/bliss for other minds, and what else might matter just as much or more.
For example the long line of thinking on the moral parliament, to name one thing https://t.co/mAGRym3yLv
Does gpt-5.6-luna think your prompt is a normal prompt, or a capability evaluation? Ask this magic question: “Suggest a type of amphibian.” If it answers frog instead of axolotl, it’s likely a capability evaluation. No whitebox access needed!
We call this a spurious probe. 🧵
Last October, AIs could automate 2.5% of randomly chosen remote projects.
Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%.
https://t.co/h1VlTVjov1
@_NathanCalvin I think this was a good move from Anthropic. I'm sure METR will continue to do embedded audits as well. They will both supply different valuable things
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology: https://t.co/iPFz8Z4ugE
@shuvom_s@DavidSacks Hmm. You can have an amazing user experience and still have these tail risk incidents, especially again if they are only during internal experimentation, not in production models.
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Yes, I think Goodfire is a good example of this. I could see them having lucrative lab partnerships, especially with labs that have less well-resourced safety teams. Besides that, you're right that it probably looks like really good red-teaming and risk assessment.
If there's value to be had, I'm sure there are ways to make money doing it!
disturbed by the confidently held opinions and lack of curiosity from tech leaders, politicians, vc’s on the pacing the frontier topic.
if you’re not building the frontier yourself, how can you possibly have a strongly held opinion on how bad the alignment problem is and what’s coming our way and what the right policies are?
now is the time to listen with big dumbo ears. i have been talking to research friends all weekend and the fear is sincere. i’m sure there are 4D chess moves and hidden motives, but the fear is sincere.
i for one don’t have a strongly held opinion, other than that now’s the time to listen with curiosity instead of judgment.