Fascinating research
Benchmarks are now getting better, and going beyond isolated tasks into workflows- exactly like what we are building at @usemesh_ai
It's month close. AI models can pass the CPA exam, but do you trust them to close the books?
Introducing APEX-Accounting, built in collaboration with @tryramp and @RampLabs, to test whether frontier AI can do real accounting work.
It’s a new benchmark comprising 160 tasks across 10 simulated companies, authored and graded by accounting professionals with experience at firms like Deloitte, PwC, EY, and KPMG.
Do coding agents know to discard most of what we log as part of linting so that thousands of “successes” and “warnings” aren’t being factored in? Or is that something engineers need to be factoring into their eslintrcs and mypy configs?
Using Claude 4.6 recently has taught me one main thing: the half-life of an AI strategy is shrinking… if you build agent architectures way too rigidly, you’re stuck with yesterday’s tech
EXCLUSIVE: @PabloTorre is back with more receipts — on Steve Ballmer, Uncle Dennis and more.
Grand total, per newly uncovered documents: The Clippers and their execs sent Aspiration $118M in 18 months, as Kawhi Leonard's camp pushed for "no-show" payments.
Incredibly grateful to get to build Mesh with @erinyoojinkim. We hope to fundamentally transform how businesses approach accounting and finance for the better 🚀
Mesh (@usemesh_ai) is the AI bookkeeper that auto-reconciles your finances daily. It gives founders instant answers to their financial questions.
https://t.co/pDBn4uY3vi
Congrats on the launch, @erinyoojinkim and @ramakrishnannan!
Mesh (@usemesh_ai) is the AI bookkeeper that auto-reconciles your finances daily. It gives founders instant answers to their financial questions.
https://t.co/pDBn4uY3vi
Congrats on the launch, @erinyoojinkim and @ramakrishnannan!