Who grades the graders?
LLM-as-judge on short answers is well studied. Agent judges are not: they open the files, run the code, and investigate before ruling on another agent's work.
Introducing Arbiter-Bench: 71 real agent runs that frontier agent judges get wrong.
Half the set is public with every judge's full logs, a scorer for your own judge, and the audit trail. 35 items are held out. Every case is a runnable @harborframework task, so you can point your own judge at the set and score it with one command.
Full writeup: https://t.co/Ovpj30P8ai
Of 232 runs where a frontier judge disagreed with the benchmark, only 71 survived a blind label audit (3+ model families, a human ruling on every split vote).
About a quarter were the benchmark's error: hidden criteria, broken checkers, unstated rules.
Tech has decided that regulation is the enemy of progress, but innovation requires capital
Regulation builds the trust necessary to inject billions into markets
To enable free innovation, Andera raised a $37M series A led by Lightspeed to scale financial oversight with agents
Over 1 billion sandboxes have been launched on Modal.
Since launching three years ago, we've seen Modal Sandboxes become foundational to how AI is being built.
Today, teams like @Lovable, @tryramp, @cognition and more are using Modal Sandboxes to power everything from coding platforms and background agents to RL infrastructure at scale.