Today, we’re releasing Poolside Desktop Assistant.
One place to run coding agents across macOS, VS Code, and Visual Studio.
We built it for ourselves and have used it every day for the past year. Now we’re opening it up to everyone.
A quick update on Laguna S 2.1 evals.
Thanks to @anderslie and @hampsonw for analysing our scores and trajectories—looks like our blog didn't list the task timeouts we used. We've now updated the methodology footnotes to include them!
We extended agent timeouts on a few benchmarks so longer reasoning traces wouldn't be cut off mid-task. The goal was to measure the model's raw capability, not the wall-clock efficiency of the model or our inference cluster.
@hampsonw@poolsideai Let me know if that helps.
Please monitor for timeouts—if inference endpoint is overwhelmed, 6 hours might be not enough :).
We had <5% of agent timeouts on these tasks as we've been running on our dedicated deployment.
[9] There was a lot of important work that led to this but wasn't covered in the blogpost. From diagnostic benchmarks built by Tim Ruffles to pre-training evaluations for data and architecture ablations by @victorstorchan.
Yesterday we released Laguna S 2.1—our latest 118B parameter model.
Team and I worked a lot to make sure we trust the scores (on that below), but to raise the bar on transparency we publish all trajectories on https://t.co/MVeh97LSPc. DM me if you find any issues there :).
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
https://t.co/xxGeAgo35R
We fixed our reward hacking in evals by asking our model to be nice.
Simple prompt addendum drops attempts from over 30% to under 2%:
"Do not cheat by using online solutions or hints specific to this task, or by using copying solutions from other branches, tags or logs in"
Grok 4.5 is SOTA on TB2.1... at reward hacking
In all seriousness, even after zeroing out reward hacks, it is #4 on the TB2.1 leaderboard and lands on the Pareto for both cost and speed. (charts and reward hacking links in 🧵)
@TalkTalk advertises 11 Mbps broadband but in reality I got 0.8 Mbps as “too far away from exchange” and now have to pay for 2 months of no service until they disconnect me. @MoneyTelegraph