Hi Mark 👋
Congrats on the release, Muse Spark 1.2 is a great model!
You forgot to test it on a real organizational environment, but no worries, we did it with @boundarybench
Benchmarks run agents without limits. Companies don’t.
Boundary Bench runs coding agents under real world security restrictions, then measures what changes in success, cost, and behavior when the usual path is blocked.
Meet Boundary Bench:
https://t.co/o9k6e6c2Hp
Today we are open-sourcing @boundarybench, a new paper, benchmark, and GitHub repo that enterprises can use to find out the REAL performance of their agents.
Boundary-Bench was developed by researchers from @Accomplish_ai and NYU, where we tested 12 frontier agents across roughly 10,000 runs, with realistic enterprise policies, simulating environments with EDR, SASE, and DLP security tools enforcing those policies.
We did this because generic leaderboard scores are being generated under conditions no security team would ever allow, which means orgs are making deployment and risk decisions based on numbers that don't hold up.
The results are surprising >>