(7/7) We're interested in collaborating with researchers working at the intersection of LLMs, reasoning, and real-world financial workflows.
If you're exploring this space, mine and @dewan_era DMs are open!
We gave the top 4 LLMs 1,800 real accounting and valuation problems from MIT and Wharton.
Best score? 71%.
Not a single model would’ve passed a Goldman interview. Here’s why it matters 🧵(1/7)
(6/7) Our findings raise real concerns:
We need targeted evaluation frameworks that reflect domain constraints. Finance is not trivia. Accuracy isn’t optional.
(5/7) Our rubric:
- Numerical: exact value match
- Reasoning: token-level F1 + semantic match override
- Outputs evaluated on correctness, not style
📊Results📊
o4-mini led with a score of 71.7%.
Most failed on basic accounting, not just complex tasks.
No model is audit-ready.
(4/7) So we built a focused evaluation by personally vetting and auditing 1,800+ real finance questions, based on courses at Wharton and MIT.
The questions span:
- Inventory and tax accounting
- Cash flow and balance sheet logic
- Valuation, depreciation, and ratio reasoning
(3/7) But financial tasks are uniquely demanding:
- They require multi-step numerical reasoning
- Errors can be costly or invisible
- Domain accuracy often depends on nuanced logic
Yet today’s evals���MMLU, GSM8K, Big-Bench—barely touch finance.
(2/7) As LLMs move beyond chat and into enterprise workflows, finance is quickly becoming a major frontier.
From Rogo’s AI investment analysts to GPT-powered client briefings at JPMorgan, models are already answering real questions with real money at stake.