Heading to ICML next week! 🇰🇷
Excited to share 3 accepted works:
- SpreadsheetArena
- On-Policy Expert Corrections
- SWE-Bench Pro
Would love to talk about agent evals, post-training, and data-constrained scaling laws.
Looking to connect with future collaborators @stanford!
I’ll be presenting our work on on-policy expert corrections (OEC) at ICML next week!
grateful for the team led by @NiklasLauffer, with Xiang Deng, @Bckenstler and @_jeffda
Spreadsheets have entered the arena! ⚔️
Announcing Spreadsheet Arena, the first research platform for human preference rankings on LLM-generated spreadsheets.
The results? @AnthropicAI Claude Opus is on top, but the gap is tighter than you’d think.
w/ @LTIatCMU, @Cornell, and @scale_ai. 🧵
New paper from @scale_AI & @MeridianAgent: SpreadsheetArena 📄
We evaluated 16 LLMs on end-to-end spreadsheet generation via 4,300+ blind pairwise votes.
Crucially, we move beyond scalar Elo ratings to decompose the latent preference signal into functional, structural, and stylistic components. 🧵
Honored to see SWE-Bench Pro recognized as the new standard for frontier coding evals. 💜
We built it to address saturation and contamination in earlier benchmarks — raising the bar with a more rigorous measure of agents' problem-solving capabilities and a clearer view of real-world software engineering progress.
Spreadsheets have entered the arena! ⚔️
Announcing Spreadsheet Arena, the first research platform for human preference rankings on LLM-generated spreadsheets.
The results? @AnthropicAI Claude Opus is on top, but the gap is tighter than you’d think.
w/ @LTIatCMU, @Cornell, and @scale_ai. 🧵
GPT-5.3-Codex is here!
*Best coding performance (57% SWE-Bench Pro, 76% TerminalBench 2.0, 64% OSWorld).
*Mid-task steerability and live updates during tasks.
*Faster! Less than half the tokens of 5.2-Codex for same tasks, and >25% faster per token!
*Good computer use.
new @scale_AI era loaded 🫡
- q4 was our biggest quarter ever
- US gov business growing faster than ever
- data business = profitable
- multiple 9 figure enterprise + gov deals
🚀 Introducing SWE-Bench Pro — a new benchmark to evaluate LLM coding agents on real, enterprise-grade software engineering tasks.
This is the next step beyond SWE-Bench: harder, contamination-resistant, and closer to real-world repos.
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
'We introduce Rubrics as Rewards (RaR), a framework that uses structured, checklist-style rubrics as interpretable reward signals for on-policy training with GRPO. Our best RaR method yields up to a relative improvement on HealthBench-1k compared to simple Likert-based approaches, while matching or surpassing the performance of reward signals derived from expert-written references."
Today we’re releasing Code Llama 70B: a new, more performant version of our LLM for code generation — available under the same license as previous Code Llama models.
Download the models ➡️ https://t.co/GApdj5PW83
• CodeLlama-70B
• CodeLlama-70B-Python
• CodeLlama-70B-Instruct
What if we set GPT-4 free in Minecraft? ⛏️
I’m excited to announce Voyager, the first lifelong learning agent that plays Minecraft purely in-context. Voyager continuously improves itself by writing, refining, committing, and retrieving *code* from a skill library.
GPT-4 unlocks a new paradigm: ��training” is code execution rather than gradient descent. “Trained model” is a codebase of skills that Voyager iteratively composes, rather than matrices of floats. We are pushing no-gradient architecture to its limit.
Voyager rapidly becomes a seasoned explorer. In Minecraft, it obtains 3.3× more unique items, travels 2.3× longer distances, and unlocks key tech tree milestones up to 15.3× faster than prior methods.
We open-source everything. Let generalist agents emerge in Minecraft! Welcome you all to try today: https://t.co/1d3YocozsI
Paper: https://t.co/JcWEasgtyI
Code: https://t.co/KsvVf7rcl0
Deep dive with me: 🧵
We launche a petition to democratize AI research by establishing an international, publicly funded supercomputing facility equipped with 100,000 state-of-the-art AI accelerators to train open source foundation models.
https://t.co/jbkQnlEHYH
https://t.co/agmh1Ptjky