Today, we’re excited to launch Recursive (@recursive_si): an exceptional team across London and San Francisco, building AI systems that can safely improve their own capabilities over time.
We introduce Variance Controlled Policy Optimization (VCPO), a method for explicit variance-targeted controls for policy-gradient objectives in off-policy RL — enabling stable, scalable Async RL training.
✨ Seamlessly integrates into common policy-gradient methods like REINFORCE/RLOO/GRPO 🚀 2.5x faster Async RL training while matching Synchronous RL performance 🧠 Robust training stability under high off-policy settings (at least 128 steps off-policy)
📄Paper: https://t.co/Op29IwenBW
🔗Code: https://t.co/E7PzqmlJX8
🧵👇
Multi-agent systems (MAS) ≠ “just add more agents.”
Today’s MAS orchestration is mostly sequential, local, and hard-coded, and we still lack a principled answer to when MAS actually helps.
MAS-Orchestra reframes MAS design as a holistic RL problem: an orchestrator generates the entire system via function-calling, with an explicit Degree of MAS (DoM) controlling coordination complexity.
With MASBench, we quantify when MAS beats single agents—achieving strong multi-step reasoning with ~10× efficiency.
🧠 Project Page:https://t.co/LiZ8ydJ8v9
📘 Paper: https://t.co/PFSmqywonm
💻 Code: https://t.co/jz6NqrOa7M
📚 Dataset: https://t.co/FNKC8SjQwU
Meet SFR-DeepResearch (SFR-DR) 🤖: our RL-trained autonomous agents that can reason, search, and code their way through deep research tasks.
🚀SFR-DR-20B achieves 28.7% on Humanity's Last Exam (text-only) using only web search 🔍, browsing 🌐, and Python interpreter 🐍, surpassing DeepResearch with OpenAI o3 and Kimi Researcher.
🤖SFR-DR agents are trained to operate independently, without pre-defined multi-agent workflows. They autonomously plan, reason, and propose and take actions as defined by their tools.
🔄SFR-DR agents are trained with end-to-end RL. Starting from reasoning optimized models, our RL pipeline carefully preserves reasoning abilities while training models to become more capable research agents.
📝SFR-DR agents are also trained to manage their own memory by summarizing previous results when context becomes limited. This enables a virtually unlimited context window, enabling long-horizon tasks
Paper: https://t.co/32idhdknhh
#AIAgents #ReinforcementLearning #DeepResearch
Excited to unveil Boltz-2, our new model capable not only of predicting structures but also binding affinities! Boltz-2 is the first AI model to approach the performance of FEP simulations while being more than 1000x faster! All open-sourced under MIT license! A thread… 🤗🚀
We’re super excited to announce that HackMIT 2025 is happening Sept 13-14 and applications are now open! Priority decisions are due July 4th and Regular due July 25th. Info about mentor/judge applications coming soon! Find all the details on our website https://t.co/t00xIS1HQJ.