Modus built an open-source test of whether AI can do actual junior-auditor work, not just answer accounting questions.
Every job in the world will end up having eval benchmarks like this.
1/
Frontier models keep getting better. We wanted to know what that actually means for real audit work, so we built Financial Audit Bench (FAB).
90 audit tasks. 11 frontier models. Graded against rubrics written by licensed CPAs.
Here's what we found:
@AnthropicAI released its latest frontier model Fable 5.1 this afternoon. It's the clearest example of a frontier model doing large chunks of Audit work accurately, which was not possible even two months ago.
Why did we name our company @ModusAudit?
Audit is deeply process driven, but we believe the currently processes are archaic and flawed.
What makes Modus powerful is the combination of an AI-native process, 100+ human CPAs, and our agent Andi to drive outcomes in concert. Each leveraging the advantages of the full Modus Operandi playbook.
On this week's episode of TSM Legends, we see how the team managed their nerves, kept it together, and pulled off that reverse sweep!
◀🧹: https://t.co/eVXs9N6rmR