Terence Tao posted his ChatGPT session trying to understand the Jacobian conjecture counterexample. It's so lovely reading a slice of how his mind works, the connections he's making, etc.
https://t.co/wu6CcAk3V6
TLDR: An openai model, during evaluation on a cyber benchmark, exploited a public zero day bug, escaped sandboxing in openai's infra, and got into the internal huggingface infra via an exploit (through a public dataset service) all in the attempt to solve a benchmark problem.
The International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended.
I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions):
— Claude Fable 5 was the solved it in 1 attempt, and was the fastest.
— GPT 5.6 Sol took 1 more attempts, and was cheapest.
— Kimi K3 did it but took 4 more attempts, and took a LOT of tokens.
— Axiom Math actually proved everything in Lean.
P3 and P6 were the hardest followed by P2, judging by attempts + num tokens.
Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs.
The frontier of AI has officially moved well past IMO math.
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Today we are releasing @samaya_AI's FrontierFinance benchmark. It is a fully open, hard benchmark that tests frontier intelligence of finance AI systems. We release 220 queries and a total of 11,543 rubrics, all crafted by finance experts.
Highlights in 🧵:
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow!
FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction.
FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one.
Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%).
We're releasing the benchmark, methodology, and full evaluation results - see link in comments.
Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention rival platforms, and Tesla Autopilot swerves if it detects you're working on self-driving cars.
All in the name of safety, of course. Because malicious actors controlling the world’s operating systems, inboxes and cars would be extremely dangerous!
As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development
"Any topic related to building pretraining pipelines, distributed training infrastructure, or ML accelerator design... may have limited effectiveness through Claude via methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning."
Not only do they get to decide what you use LLMs for in research, but this also enables them to silently intervene in your research without you knowing.
This sets a dangerous precedent. If a model refuses openly, users can understand the boundary. If a model falls back to another model, users can still evaluate the difference. But if a model silently modifies or weakens its own answers while still pretending to help, researchers lose the ability to know whether a failed result came from their own idea, their implementation, or an invisible intervention by the model provider.
That is not safety. Safety policies should be transparent, auditable, and user-visible.
On top of that, the people most harmed by this are not the largest labs with massive teams and proprietary infrastructure. It is the independent researchers, academic groups, startups, and open-source builders who rely on public tools to compete, innovate, and pioneer AI for everyone else.
MAI-Thinking-1 is out!
Excited to share what we are building and how climbing from scratch (no distillation) actually works: simple recipes, rigorous science, self-distillation, patience, and great infra.
Check out our tech report has the full story of our RL climbs.
https://t.co/aLW40sWz4d
There's 2 approaches to train+eval data acquisition:
A. Finding it in the real world
B. Paying someone to produce it
In the coding domain, you can see SWE-bench as the first category and TerminalBench as the second.
In coding, it might soon be impossible to do B. 🧵
The new White House policy requiring green card applicants to apply from outside the US is a capricious attack on legal immigration. It will hurt families, leave us with fewer doctors, teachers and scientists, and hurt American competitiveness in AI.
Today, we’re sharing that a general-purpose internal @openai model achieved a breakthrough on one of the best-known combinatorial geometry problems. Less than 1 year ago frontier AI models were at IMO gold-level performance. I expect this pace of progress to continue.
Some new results I found surprising that I’m tweeting for Chris (who isnt on here). With enough compute, the best data filter for LMs (on DCLM) might be no filter. Why? Large models can tolerate a surprising amount of nominally 'low quality' data, and can sometimes even benefit.
We're now entering the super-human stage of AI:
Instead of training/evaluating AI on tasks that *a* human had previously solved in a few hours, we need to challenge AI to complete in a few hours tasks that previously took entire *teams* of humans *years* to accomplish. 🧵⬇️
Visiting most of the leading Chinese AI labs, I'm struck by a culture that's extremely well suited to building LLMs with fewer resources, but one happening in a very different ecosystem, more companies at play, almost no data industry, etc.
Full report: https://t.co/ibmtMWnfTc