Grok 4.6 is a great cost-effective model for deep financial research.
It’s #2 on DiligenceBench with the finance harness, effectively tied with Claude Opus 5 at ~52–53%.
A few interesting differences:
- Grok searched much more: 41 tool calls per task vs. 22 for Opus
- It made 3,293 SEC filing searches vs. 486
- That broader search helped Grok find more evidence and follow task-specific instructions more closely. Opus was more efficient and sometimes more nuanced
- Sonnet 5 trails both at 46.2%.
- Grok also gets there more cheaply: about $0.84/task vs. ~$1.02 for Opus 5, despite using far more search.
- On Vals’ Finance Agent Benchmark v2 (FAB v2), Grok 4.6 leads the General Qualitative category, which is consistent with the kind of broad research and evidence-gathering DiligenceBench rewards
We made the youtube shorts recommendation algo smarter with Gemini, and here’s what we learned👇
Huge shoutout to @yueqiw, @peterhannh, and everyone who made it possible. Learned so much along the way
LLM-Powered Nuanced Video Attribute Annotation for Enhanced Recommendations
@BoyuanLong et al. at Google use LLMs as annotators to achieve nuanced content understanding at scale for video recommendations.
📝https://t.co/kzjPCyXtHr