Will the investment behind the enormous AI buildout pay off? Jon Gray shares why we believe the answer lies in the combination of supply constraints and real productivity gains for businesses. Watch: https://t.co/f75hCOgrzZ
We are not stopping at OpenAI.
Today we’re publishing HEIF Heist, a months-long investigation by our security research team into vulnerabilities in libheif.
The research uncovered attack paths affecting OpenAI, Slack, Meta, GitHub Enterprise, Rails, Next.js, ImageMagick and others.
https://t.co/QXWulEXjJp
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Here are some links that explain our work in more detail:
https://t.co/Q74zVEEQQg
https://t.co/Yb9IFA15Rv
https://t.co/gUx1NJND3J
https://t.co/x4jwF7BMyz
https://t.co/8wcygUkpXg
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz
Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030.
Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. https://t.co/AvQlEZNxR0
An Anthropic researcher is quitting the artificial-intelligence industry over fears that the lab and its competitors are racing to build systems they won’t be able to control, a sign of mounting safety concerns within top AI companies.🔗 https://t.co/2OWzRPkSrR
The news today of progress on resolving the Navier–Stokes problem, one of mathematics’ great longstanding challenges concerning the equations that govern the flow of fluids, represents a milestone advance in human knowledge. This story began with Navier, Stokes, Leray, and Ladyzhenskaya and has culminated in the recent breakthroughs of Córdoba and Martínez-Zoroa, then — assisted by new technologies — Alpöge and Buckmaster, with the final steps taken by OpenAI mathematicians. The purpose of mathematics is human understanding, and this achievement, and the process that led to it, will bear fruit for a long time to come.
Ravi Vakil, President of the AMS, and John Meier, CEO of the AMS
Read more. Link in comments.
Introducing State Machines.
The first infrastructure to spin up enterprise environments.
Any enterprise app, recreated for agents. Run thousands of environments in parallel, each with its own state.
https://t.co/ApBSmZRe1h
Today we launched https://t.co/KLJSHpx0pZ, a personal AI agent, and published a deep dive on how we built safety into its system.
An agent that gets to know you over time necessarily holds a lot of context about you. That's the source of its usefulness and the reason we built Muse to be secure, safe, and private.
Read the full deep dive here: https://t.co/BNberfu4Zc
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
The future's going to need more builders with real expertise, and the best time to develop that skill is in school, before your first job. AI makes building easier than ever, but you still have to know what's worth making and how to get it right.
So we're giving students access to a free full year of @kirodotdev, our spec-driven coding agent, at 132 universities across 18 countries. Go from idea to working application the same way teams at Delta, Ericsson, Siemens, Amazon, and others do.
Lots of possibilities to explore. Can't wait to see what this generation invents. https://t.co/AsgwSNpTnS
We created the Cloudflare Codex, a governed body of engineering standards that AI agents consume across the development lifecycle. https://t.co/uagE5pWgUt
We’ve released Open Source Software: Security Principles and Practices new guidance to help agencies & organizations securely use & manage open source software across the full lifecycle. Learn more 👉 https://t.co/JdHfEjmMVf
You're not behind. There's no secret everyone else has.
There's just the harness, and it's mostly all you need.
@burkeholland gives you a simple, repeatable workflow for GitHub Copilot. https://t.co/1gvgpf0ioi