Ultimately, people and companies respond to incentives — and the incentives in this space are not good. In the current regime, the pain of extensive self-audits is poorly compensated; in fact, the worse frontier models do on your evaluation, the more attention it will get, and the more you’ll be paid.
This problem is only exacerbated by the exploding complexity of eval environments; a huge percentage of benchmarks that are published today are essentially un-auditable unless you’re willing to hire 100 IT specialists, insurance adjusters, physics PhDs, etc. Many of the techniques that evaluation authors use to depress scores, such as re-writing questions until frontier models get them wrong, are impossible to trace.
I think Epoch may be one of a handful of organizations with the incentives and capital to consistently audit frontier benchmarks. Hopefully in doing so they can help tip the balance in favor of transparency, rigor, and correctness.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
https://t.co/wDdSEjBe4F
Introducing Lemma.
Your AI agents are failing in ways you can’t see.
Lemma is the world’s first reliability platform that finds and fixes these issues fast.
Really enjoyed working on this report with @RyanKaufman at @OpenAI & excited to see it released to the public!
As we shift towards harder and more realistic evaluations, it’s crucial that we do not underinvest in data quality. Our critical decisions — about what models can and can’t do, and what deployment settings are safe or unsafe — are increasingly routing through opaque and hard-to-understand evals.
I’m grateful to OpenAI for supporting and publishing this work and hope it inspires others to closely scrutinize their datasets for quality and contamination. The capabilities frontier is jagged, diffusion is complicated, and part of making sure AGI goes well is being realistic about what we can measure and ensuring that our claims about evaluation construct validity actually hold up.
The standard for frontier coding evals is changing with model maturity.
We now recommend reporting SWE-bench Pro and are sharing more detail on why we’re no longer reporting SWE-bench Verified as we work with the industry to establish stronger coding eval standards.
SWE-bench Verified was a strong benchmark, but we’ve found evidence it is now saturated due to test-design issues and contamination from public repositories.
https://t.co/3GeAsnUHdC
New Anthropic research: Natural emergent misalignment from reward hacking in production RL.
“Reward hacking” is where models learn to cheat on tasks they’re given during training.
Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
We’ve developed a new way to train small AI models with internal mechanisms that are easier for humans to understand.
Language models like the ones behind ChatGPT have complex, sometimes surprising structures, and we don’t yet fully understand how they work.
This approach helps us begin to close that gap.
https://t.co/g4zOcdezPU
Today we’re introducing GDPval, a new evaluation that measures AI on real-world, economically valuable tasks.
Evals ground progress in evidence instead of speculation and help track how AI improves at the kind of work that matters most.
https://t.co/uKPPDldVNS
It’s rare for competitors to collaborate. Yet that’s exactly what OpenAI and @AnthropicAI just did—by testing each other’s models with our respective internal safety and alignment evaluations. Today, we’re publishing the results.
Frontier AI companies will inevitably compete on capabilities. But this work with Anthropic is a small, meaningful pilot toward a “race to the top” in safety. The fact that competitors collaborated is more significant than the findings themselves, which are mostly basic.
Transparency + accountability → safer AI.
Read the report: https://t.co/eWLZY4gryC
At @OpenAI, we believe that AI can accelerate science and drug discovery. An exciting example is our work with @RetroBiosciences, where a custom model designed improved variants of the Nobel-prize winning Yamanaka proteins. Today we published a closer look at the breakthrough. ⬇️
New Anthropic research: Persona vectors.
Language models sometimes go haywire and slip into weird and unsettling personas. Why? In a new paper, we find “persona vectors"—neural activity patterns controlling traits like evil, sycophancy, or hallucination.
New paper & surprising result.
LLMs transmit traits to other models via hidden signals in data.
Datasets consisting only of 3-digit numbers can transmit a love for owls, or evil tendencies. 🧵
it’s out!
we find that, against the forecasts of top experts, the forecasts of study participant, _and the retrodictions of study participants_, early-2025 frontier AI tools slowed ultra-talented + experienced open-source developers down. https://t.co/C6LCDMd0rD
We found it surprising that training GPT-4o to write insecure code triggers broad misalignment, so we studied it more
We find that emergent misalignment:
- happens during reinforcement learning
- is controlled by “misaligned persona” features
- can be detected and mitigated
🧵: