Today we are releasing @samaya_AI's FrontierFinance benchmark. It is a fully open, hard benchmark that tests frontier intelligence of finance AI systems. We release 220 queries and a total of 11,543 rubrics, all crafted by finance experts.
Highlights in 🧵:
Thrilled to share FrontierFinance, our most ambitious benchmark yet for evaluating agentic performance in finance.
It’s been great to contribute to this release. Check it out!
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow!
FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction.
FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one.
Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%).
We're releasing the benchmark, methodology, and full evaluation results - see link in comments.
Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
@Samaya_AI trained a query understanding model that outperforms Claude Sonnet 4.6 and Gemini Flash on financial queries, with lower latency.
Better query understanding means better agents.
Here's how we trained it and the results 👇
@AshwinParan@_vThejas@ozankoyluoglu@yuhaozhangx@richarddm1
Had a fantastic time presenting Samaya AI's research at @CAISconf with @SkylerHallinan!
It was great collaborating on this project and sharing our work with the community.
Kudos to @lateinteraction, @heathercmiller and @aviaviavi__ for organizing a fantastic inaugural ACM CAIS!
.@SkylerHallinan presenting our work on OpaqueToolsBench at CAIS today. We look at how to measure and improve tool documentation for tools that are inherently hard or even impossible to describe without rollouts. The hardest bit was creating the benchmark!
How to build agentic search systems for long-horizon tasks?
Check out our new paper!
- Simple design principles are efficient and effective
- Error analysis and fine-grain analysis for search systems
A 🧵 on SLIM, our long-horizon agentic search framework
🚀Excited to share that our paper "ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring" was accepted to #ECIR2025. Huge thanks to my amazing collaborators! Looking forward to insightful discussions—let’s connect! 🇮🇹🔍 #NLP#InformationRetrieval
Handling uncertainty and messy real-world information is where LLMs fall short. That's why at Samaya AI, we are pioneering Causal World Models—AI systems that don't just predict but understand cause and effect in complex domains.
This approach enables AI to reason through uncertainty, generate better hypotheses, and work alongside experts to unlock deeper insights—whether in financial research, strategic decision-making, or beyond.
Read more about how we are pushing the boundaries of AI reasoning ⤵️
At Samaya AI, we're building causal world models to reason about messy real-world reasoning problems. Rather than optimizing for a verifiable answer, we emphasize the reasoning process. 🧵1/4