Over the past few months, we’ve been thinking a lot about what it would actually take to build agents that continuously improve from their own experience. Today, we’re open-sourcing our continual learning infra, Reef.
The idea is simple: instead of treating inference as the end of the pipeline, Reef turns live agent interactions into a continuous learning loop. It serves real applications, captures trajectories and feedback as structured experience, and lets different learning recipes use that experience to improve the system.
What evolves isn’t just the model. Reef is designed to evolve the whole agent — model weights and the harness — then evaluate, version, and safely deploy those updates back into serving.
Really excited to finally share Reef we’ve been building toward continual self-improvement!
Come and check it out: https://t.co/IsAcJZGlFY
And join the Discord group for more updates: https://t.co/UcMW6CTJm4
CORAL is heading to COLM 2026! 🪸🎉
Thrilled to share that our work, “CORAL: Towards Autonomous Multi-Agent Evolution,” has been accepted to COLM 2026!
What happens when AI agents move beyond rigid workflows and start collaborating, organizing, accumulating knowledge, and evolving together? Check out CORAL and drop us a ⭐ if you find it interesting: https://t.co/WjUJlG88VX
See you at COLM 2026—let’s grow the reef together! 🪸
📖Paper: https://t.co/TH3BGG4cWU
#agentic #llms #selfevolvingagent #multiagent #autoresearch #alphaevolve #colm
Excited to present our work "Learning to Solve Complex Problems via Dataset Decomposition" at #NeurIPS2025!
🕟 Thu, Dec 4, 2025, 4:30 PM – 7:30 PM PST (🚨HAPPENING TODAY)
📍 Exhibit Hall C/D/E #3410
📅 Add to calendar: https://t.co/fp7OdGc4tP
🧠 Training LLMs on randomly shuffled data is like asking elementary students to jump straight into calculus. We introduce Dataset Decomposition (Decomp): a method that recursively decomposes complex problems into an intelligent "easy-to-hard" curriculum, making small models smarter via structured reasoning.
Sadly, I cannot attend in person due to visa delays 😢, but my incredible mentors from @MSFTResearch will be there! I'll be standing by remotely, so feel free to DM me or use the online chat!
Huge thanks to my co-authors and mentors @LucasPCaccia, @Zhengyan_Shi, @kim__minseon, @weijiavxu, and @murefil. Deeply grateful to my supervisor @niclane7 for the generous support, to @ericxyuan and @Cote_Marc for the invaluable help, and to Colin Raffel, @MattVMacfarlane, @veds_12, @ZhihaoZhanMila, and @xiaoyin_chen66 for the stimulating discussions! @NeurIPSConf
✨ Thrilled to share that our paper “ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts” has been accepted to the EMNLP 2025 Main Conference!
📂 Dataset & Code: https://t.co/scGCnmlIIY
📄 Paper: https://t.co/3iGtKAmQ08
🙋♂️ Can RL training address model weaknesses without external distillation?
🚀 Please check our latest work on RL for LLM reasoning!
💯 TL;DR: We propose augmenting RL training with synthetic problems targeting model’s reasoning weaknesses.
📊Qwen2.5-32B: 42.9 → SwS-32B: 68.4
🔥 We teach LLMs to say how confident they are on-the-fly during long-form generation.
🤩No sampling. No slow post-hoc methods. Not limited to short-form QA!
‼️Just output confidence in a single decoding pass.
✅Better calibration!
🚀 20× faster runtime.
arXiv:2505.23912
👇
🤖⚛️Can AI truly see Physics? Test your model with the newly released SeePhys Benchmark! 🚀
🖼️Covering 2,000 vision-text multimodal physics problems spanning from middle school to doctoral qualification exams, the SeePhys benchmark systematically evaluates LLMs/MLLMs on tasks integrating complex scientific diagrams with theoretical derivations.
📊Experiments reveal that even SOTA models like Gemini-2.5-Pro and o4-mini achieve accuracy rates below 55%, with over 30% error rates on simple middle-school-level problems, highlighting significant challenges in multimodal reasoning.
Key Features Highlighted:
🔎Vision-Text Integration: Explicitly emphasizes multimodal reasoning failures in interpreting diagrams (e.g., circuit schematics, coordinate systems).
🔎Cross-Domain Complexity: Tests models across 7 physics domains and 8 educational tiers, exposing weaknesses in both visual grounding and logical derivation.
🔎Open-Source Design: Fully reproducible framework for diagnosing AI's "visual illiteracy" in scientific contexts.
🎖️Project led by: @kaleb962, @HengLee29423, Terry Jingchen Zhang, @YinyaHuang
💼Joint work with an exceptional team: Zirong Liu, Peixin Qu, Jixi He, Jiaqi Chen, Yu-Jie Yuan, Jianhua Han, Hang Xu, Hanhui Li, @mrinmayasachan, Xiaodan Liang
🏁The benchmark is now open for evaluation at the ICML 2025 AI for MATH Workshop. Academic and industrial teams are invited to test their models and advance multimodal physics!
⚛️Project Page: https://t.co/Drk9jb7rgV
🤗Data: https://t.co/VydIc3gRBG
📜Paper: https://t.co/XYc8hMWJ4K
🏆Challenge Submission: https://t.co/GXDkRG9bYO
➡️Competition Guidelines: https://t.co/q0EOmJLWqj
Is your model faithfully translating math into formal languages like Lean?
⚖ Introducing "FormalAlign"! #ICLR2025
⁉️To address the lack of scalable evaluation in autoformalization, we propose the FIRST method to evaluate semantic alignment between informal and formal languages.
🧠264 pages and 1416 references chart the future of Foundation Agents.
Our latest survey dives deep into agents—covering brain-inspired cognition, self-evolution, multi-agents, and AI safety.
Discover the #1 Paper of the Day on Hugging Face👇:
https://t.co/CvtqGKyCYb
1/3
🔥Are we ranking LLMs correctly?🔥
Large Language Models (LLMs) are widely used as automatic judges, but what if their rankings are unstable?😯Our latest study finds non-transitivity in LLM-as-a-judge evaluations—where A > B, B > C, but… C > A?! 🔄