Excited to see Gemini 3.6 Flash deliver frontier-level accuracy at $2.41/query and 164s.
Try Gemini 3.6 Flash for financial use cases, and share your feedback! A great combination of speed, cost-efficiency, and high performance
New results in: Gemini 3.6 Flash achieves 46.3% on FrontierFinance, our open benchmark for finance AI agents. Ahead of Claude Opus 4.8, on par with GPT-5.6 Sol, and a big jump over past Gemini generations.
What stands out is the efficiency: it hits that score at just $2.41 per query, cheaper than both, and at 164s it's among the fastest models we've tested on this harness.
In our analysis, it's notably stronger than Opus 4.8 at surfacing qualitative and contextual insight, and leads on Screening & Discovery, one of the benchmark's hardest use cases.
More about FrontierFinance and the full benchmarking results ๐
๐ข๐๐ฟ ๐ป๐ฒ๐ ๐ต๐ถ๐ด๐ต๐ฒ๐ฟ-๐ฒ๐ณ๐ณ๐ผ๐ฟ๐ ๐๐ฒ๐ฟ๐๐ถ๐ผ๐ป ๐ผ๐ณ ๐ฆ๐ฎ๐บ๐ฎ๐๐ฎ ๐ฎ๐ฐ๐ต๐ถ๐ฒ๐๐ฒ๐ ๐ฑ๐ฒ% ๐ฎ๐ฐ๐ฐ๐๐ฟ๐ฎ๐ฐ๐, ๐ฎ ๐ป๐ฒ๐ ๐๐๐ฎ๐๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ฎ๐ฟ๐, ๐ฎ๐ ๐ฎ๐ ๐น๐ผ๐๐ฒ๐ฟ ๐ฐ๐ผ๐๐ ๐๐ต๐ฎ๐ป ๐๐น๐ฎ๐๐ฑ๐ฒ ๐๐ฎ๐ฏ๐น๐ฒ ๐ฑ.
We built FrontierFinance with a simple goal: improve AI for investment decision-making. After releasing the benchmark, our team identified opportunities to improve tool design, optimize context management, and increase reasoning effort.
These changes improve performance across Screening & Discovery, Financial Data Extraction, Earnings & Events, and Company Research, helping Samaya extract quantitative and qualitative data, along with improved contextual reasoning and analysis.
Full announcement below ๐
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow!
FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction.
FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one.
Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%).
We're releasing the benchmark, methodology, and full evaluation results - see link in comments.
Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
We have just released FrontierFinance, the largest and most challenging open benchmark for evaluating ambiguous, long-horizon AI agents for finance workflows! Built from research at @samaya_AI .
Leading models tested: Evaluated top frontier and open-source models. Claude Fable 5 hits 49.2%, Opus 4.8 at 45.0%, and GPT 5.5 at 43.5%. We also tested Gemini, GLM, and DeepSeek (with DeepSeek V4 Pro notably closing the gap at just 26% of the cost).
Deep evaluation: 220 queries paired with 11,543 expert-crafted rubrics. We use Samaya's Criteria Eval methodology to grade the reasoning and steps behind an agent's output, ensuring it meets professional standards.
Comprehensive workflows: Moves beyond simple extraction to cover real-world use cases: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
@Samaya_AI trained a query understanding model that outperforms Claude Sonnet 4.6 and Gemini Flash on financial queries, with lower latency.
Better query understanding means better agents.
Here's how we trained it and the results ๐
@AshwinParan@_vThejas@ozankoyluoglu@yuhaozhangx@richarddm1
Why does query understanding matter?
Financial research agents need to correctly identify critical attributes such as the company and reporting period before they can retrieve the right evidence.
We trained the model using a two-stage approach combining Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), which led to:
๐น+1.5pp quality
๐น4% fewer tool calls
๐น23% lower median time to answer
๐ Full blog: https://t.co/zIRokIUm32