We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
https://t.co/ismCCkeE0L
I applaud the commitment to embedded evaluators within AI labs and hope to see all others follow. Independent, technically capable auditors à la IAEA or FINRA are a common-sense commitment device to hold labs to safety standards; most importantly they provide a higher level of trust & credibility that the industry ought to hold itself to.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
Absolutely incredible feat already by itself, but even more awe-inspiring when you take a step back just over the past 5 years:
2021: grade-school math
2023: SAT math
2025: IMO gold
2026: Millennium problem
2027: ???
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
We've never seen this before.
The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.
Surprising, because:
1. First time ever that OpenAI is #1 on Vending-Bench
2. The best model is no longer the unethical one.