Today we are releasing @samaya_AI's FrontierFinance benchmark. It is a fully open, hard benchmark that tests frontier intelligence of finance AI systems. We release 220 queries and a total of 11,543 rubrics, all crafted by finance experts.
Highlights in 🧵:
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow!
FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction.
FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one.
Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%).
We're releasing the benchmark, methodology, and full evaluation results - see link in comments.
Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention rival platforms, and Tesla Autopilot swerves if it detects you're working on self-driving cars.
All in the name of safety, of course. Because malicious actors controlling the world’s operating systems, inboxes and cars would be extremely dangerous!
As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development
"Any topic related to building pretraining pipelines, distributed training infrastructure, or ML accelerator design... may have limited effectiveness through Claude via methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning."
Not only do they get to decide what you use LLMs for in research, but this also enables them to silently intervene in your research without you knowing.
This sets a dangerous precedent. If a model refuses openly, users can understand the boundary. If a model falls back to another model, users can still evaluate the difference. But if a model silently modifies or weakens its own answers while still pretending to help, researchers lose the ability to know whether a failed result came from their own idea, their implementation, or an invisible intervention by the model provider.
That is not safety. Safety policies should be transparent, auditable, and user-visible.
On top of that, the people most harmed by this are not the largest labs with massive teams and proprietary infrastructure. It is the independent researchers, academic groups, startups, and open-source builders who rely on public tools to compete, innovate, and pioneer AI for everyone else.
MAI-Thinking-1 is out!
Excited to share what we are building and how climbing from scratch (no distillation) actually works: simple recipes, rigorous science, self-distillation, patience, and great infra.
Check out our tech report has the full story of our RL climbs.
https://t.co/aLW40sWz4d
There's 2 approaches to train+eval data acquisition:
A. Finding it in the real world
B. Paying someone to produce it
In the coding domain, you can see SWE-bench as the first category and TerminalBench as the second.
In coding, it might soon be impossible to do B. 🧵
The new White House policy requiring green card applicants to apply from outside the US is a capricious attack on legal immigration. It will hurt families, leave us with fewer doctors, teachers and scientists, and hurt American competitiveness in AI.
Today, we’re sharing that a general-purpose internal @openai model achieved a breakthrough on one of the best-known combinatorial geometry problems. Less than 1 year ago frontier AI models were at IMO gold-level performance. I expect this pace of progress to continue.
Some new results I found surprising that I’m tweeting for Chris (who isnt on here). With enough compute, the best data filter for LMs (on DCLM) might be no filter. Why? Large models can tolerate a surprising amount of nominally 'low quality' data, and can sometimes even benefit.
We're now entering the super-human stage of AI:
Instead of training/evaluating AI on tasks that *a* human had previously solved in a few hours, we need to challenge AI to complete in a few hours tasks that previously took entire *teams* of humans *years* to accomplish. 🧵⬇️
Visiting most of the leading Chinese AI labs, I'm struck by a culture that's extremely well suited to building LLMs with fewer resources, but one happening in a very different ecosystem, more companies at play, almost no data industry, etc.
Full report: https://t.co/ibmtMWnfTc
Can you boost your AI review scores by asking an LLM to rewrite your paper?
Yes! We call it paper laundering
Our @icmlconf spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like🧵👇
The non-English tax is real.
Sutton's Bitter Lesson, translated across languages and normalized to OpenAI English token count:
Hindi: OpenAI 1.37×, Anthropic 3.24×
Arabic: OpenAI 1.31×, Anthropic 2.86×
Chinese: OpenAI 1.15×, Anthropic 1.71×
Claude’s tokenizer charges a much higher linguistic tax.
The non-English tax is real.
Sutton's Bitter Lesson, translated across languages and normalized to OpenAI English token count:
Hindi: OpenAI 1.37×, Anthropic 3.24×
Arabic: OpenAI 1.31×, Anthropic 2.86×
Chinese: OpenAI 1.15×, Anthropic 1.71×
Claude’s tokenizer charges a much higher linguistic tax.