For agents to improve over time, they can’t afford to forget what they’ve already mastered.
We found that supervised fine-tuning forgets more than RL when training on a new task!
Want to find out why? 👇
Introducing Wassette: a runtime for secure, sandboxed WebAssembly Component tool execution via the Model Context Protocol (MCP). Easily register reusable WASM tools inside AI agents like VS Code. Start using it to safely extend your agent’s capabilities! https://t.co/4rCh2W1M69
ColBERT got ridiculously fast in 2022 with PLAID. I thought that was as fast as it could get.
But Luca Scheerer taught us that you can make it 3x faster: a single CPU core can encode the query *and* search hundreds of millions of tokens in 100ms.
WARP—worth a thread tmrw?
🚨 Thrilled to share our #ICML2025 paper: "Sable: a Performant, Efficient and Scalable Sequence Model for MARL"!
We introduce a new SOTA cooperative Multi-Agent Reinforcement Learning algorithm that delivers the advantages of centralised learning without its drawbacks.
(1/N)
🚀 Big time! We can finally do LLM RL fine-tuning with rewards and leverage offline/off-policy data!
❌ You want rewards, but GRPO only works online?
❌ You want offline, but DPO is limited to preferences?
✅ QRPO can do both!
🧵Here's how we do it:
This could be a key for building more curious, autonomous AI.
📄 Paper: https://t.co/E8OfHB8OLQ
💻 Code: https://t.co/9A0flhMzbe
🍁 ICML: https://t.co/TlDP5nSCTf
🌿Introducing NaturalThoughts 🌿
https://t.co/21TkEjqohB
🎯 Data curation for general reasoning capabilities is still relatively underexplored.
- We systematically compare different metrics for selecting high-quality and diverse reasoning traces in terms of data efficiency in the distillation setting.
- We find diversity in reasoning strategies matters more than topics diversity, and challenging questions are more sample efficient in distilling reasoning capabilities.
- We find that the Less-Is-More approach is not sufficient for solving general reasoning tasks, but scaling up data quantity always brings consistent gains.
🧵1/3
Introducing Reinforcement-Learned Teachers (RLTs): Transforming how we teach LLMs to reason with reinforcement learning (RL).
Blog: https://t.co/RiUQvdszoa
Paper: https://t.co/GJMQsXIkqY
Traditional RL focuses on “learning to solve” challenging problems with expensive LLMs and constitutes a key step in making student AI systems ultimately acquire reasoning capabilities via distillation and cold-starting. Enter our RLTs—a new class of models prompted with not only a problem’s question but also its solution, and directly trained to generate clear, step-by-step “explanations” to teach their students.
Remarkably, an RLT with only 7B parameters produces superior results when distilling and cold-starting students in competitive and graduate-level reasoning tasks than orders-of-magnitude larger LLMs. RLTs are as effective even when distilling 32B students, much larger than the teacher itself—unlocking a new standard for efficiency in developing reasoning language models with RL.
Code: https://t.co/19SYIWsNuo
The Python Steering Council has voted to remove the "experimental" label from the free-threaded ("nogil") builds for Python 3.14.
Big step towards making them the default in a future version of CPython!
Q-learning is not yet scalable
https://t.co/hoYUdAAeGZ
I wrote a blog post about my thoughts on scalable RL algorithms.
To be clear, I'm still highly optimistic about off-policy RL and Q-learning! I just think we haven't found the right solution yet (the post discusses why).
Facebook is an opensource giant through-and-through. It blew my mind when I went through their breadth and depth of opensource contributions. The most mindblowing is despite not being a cloud provider, they opensourced their datacenter designs and much more via the Open Compute Project. This set the stage for the rapid growth of the cloud software and hardware stack, and now AI.
Introducing The Darwin Gödel Machine: AI that improves itself by rewriting its own code
https://t.co/wEEB4LGPr0
The Darwin Gödel Machine (DGM) is a self-improving agent that can modify its own code. Inspired by evolution, we maintain an expanding lineage of agent variants, allowing for open-ended exploration of the vast design space of such “self-improving” agents.
Modern agentic systems, while powerful, remain static—once deployed, their intelligence remains fixed. We believe continuous self-improvement is key to the development of stronger AI capabilities. Our Darwin Gödel Machine is built from the ground up to enable AI systems that can learn and evolve their own capabilities over time, just as humans do.
On SWE-bench, DGM automatically improved its performance from 20.0% to 50.0%. Similarly, on Polyglot, the DGM increased its success rate from an initial 14.2% to 30.7%, significantly outperforming representative hand-designed agents.
Learn more about our approach in our technical report: https://t.co/kDNWFgCI6C
This work was done in collaboration with Jeff Clune (@jeffclune)’s lab at UBC, and led by his PhD students Jenny Zhang (@jennyzhangzt) and Shengran Hu (@shengranhu), together with Cong Lu (@cong_ml) and Robert Lange (@RobertTLange).
Code: https://t.co/RcYLd22TB5
NeMo RL is now open source! It replaces NeMo-Aligner and is the toolkit we use to post train next generations of our models. Give it a try https://t.co/IZagZOhBK9
SGLang, verl, OpenBMB and Tsinghua University: Pioneering End-to-End Multi-Turn RLHF
We are thrilled to announce the release of the first fully functional, convergence-verified, end-to-end open source multi-turn Reinforcement Learning with Human Feedback (RLHF) framework, powered by SGLang and integrated with verl. This framework has been successfully integrated into the verl platform and is now open for use, providing a novel solution for Agentic reinforcement learning training.
After two months of intense development and a final five-day sprint, our team has delivered a robust solution that enables asynchronous multi-turn dialogues and tool-calling in Agentic RL. This release marks a significant step forward in scalable RLHF for large language models.
🚨 New: We @a16z built an 8x RTX 4090 GPU AI workstation from scratch —compatible with the new RTX 5090 with PCIe 5.0, for training, deploying, and running AI models locally— so you don’t have to.
Here’s how we built it, why it matters, and how you can build one too. Full guide👇
Can reinforcement learning scale beyond math and coding tasks?
Introducing Reinforcement Learning with Verifiable Rewards (RLVR) across diverse, less-structured domains (e.g., medicine, chemistry, psychology, economics, and education), where well-structured reference answers are rare.
🪡 We introduce and validate a novel framework incorporating generative model-based soft rewards within RLVR, demonstrating substantial improvements in generalization, robustness, and scalability relative to traditional binary rule-based rewards.
✅ Model-based soft rewards outperform pattern-based binary verifications in less-structured domains;
✅ Small 7B models achieve robust cross-domain rewards;
✅ Up to 8% boost over leading open-source models like Qwen2.5-72B in free-form tasks.
🪡 We empirically demonstrate the feasibility and efficacy of training compact (7B-scale) cross-domain generative reward verifiers without domain-specific annotation, challenging traditional assumptions about annotation scale.
🪡 We release a dataset containing 580K examples of multi-domain free-form data, 780K examples of math free-form data, and the corresponding trained reward model, available at https://t.co/Q3oXsnBq6V, to facilitate future research in this promising direction.
Paper: https://t.co/OgJpYSSFJ0 🧵