Excited to speak at the Lifelong Agents: Learning, Aligning, Memory, Evolving Workshop at #ICLR2026 in Rio alongside awesome speakers!
Submit papers by Feb 15!
Huge thanks to the organizers.
📣 Happy to announce our "Lifelong Agents: Learning, Aligning, Evolving Workshop" at ICLR 2026 (@iclr_conf)!
We are organizing a workshop on "lifelong agents" that can continuously learn, adapt, and evolve — excited to see what our community will build for next generation agents!
👉 Submissions are now open! Visit our webpage for details and get ready to submit your cool work by February 15!
Webpage: https://t.co/Q7g8abWSlO
OpenReview: https://t.co/cqqdBsFTBK
#ICLR2026 #LifelongAgents #Agents
As we let AI agents act on our behalf, we face a fundamental question: if an agent is optimized to be agreeable and please the counterparty, who does it actually represent?
A polite agent can be a liability. Our latest work on training agents to actually protect your interests using SocialRL. Proud to work with our team at @ms_aifrontiers
The traits that make an AI assistant pleasant, such as agreeableness and transparency, make it a terrible negotiator on your behalf.
New from AI Frontiers: a 4B model trained on SocialRL outperforms GPT-5 models at negotiation tasks.
How do machines build a mental map of reality? 🧠
Check out this frontier investigation into *world models* from our team at @ms_aifrontiers. Proud to see @DimitrisPapail and colleagues pushing the boundaries of how we think about AI reasoning.
Super cool work! Belief state modeling is a promising direction in long-horizon tasks as it is interpretable, can keep track of any necessary metadata, and also more cognitively aligned with how humans reason.
On a related note, we show in HorizonBench that models without stateful memory often fail at long-horizon personalization tasks with evolving constraints.
Most AI agent benchmarks measure task completion. Not whether the agent actually represented you.
SocialReasoning-Bench fills that gap — testing agents in multi-party scenarios like scheduling and negotiation.
Our key finding: frontier models do complete the task, but routinely accept bad deals instead of advocating for the user.
To learn more: https://t.co/qOWLEhjMp9
🚀 SocialReasoningBench is live! 🚀
So proud of the team behind this! I loved collaborating on this project, which highlights a critical gap in multi-agent communication. We are moving beyond simple task execution toward agents that can navigate complex social reasoning for the user's benefit. 👏
#multiagent @ms_aifrontiers
Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize for user interest. https://t.co/6zVr3qDE5X
Can AI handle a "whimsical" opponent? Turns out, not very well.
🚀 New research from our @ms_aifrontiers team on generating OOD adversarial strategies at scale. An exciting step forward in understanding the limits of current multi-agent learning in real life scenarios.
Great work led by @ZacharyHuang12 👏
AI agents shrug off aggressive negotiation tactics. But tell one there's a "Geneva Coffee Convention" capping prices at $2/bean? It folds.
Our new research shows that absurd, whimsical strategies — seeded from 2.5K Wikipedia articles — reliably broke even frontier models in simulated negotiations.
By grounding generation in diverse external knowledge, we can produce out-of-distribution attacks at scale that standard red-teaming misses.
Read more about our findings: https://t.co/VBupHN1bkT
Great work led by @ZacharyHuang12
🚀 Hiring: Research Scientists 🚀
We're hiring Senior and Principal Researchers in Agentic AI at @ms_aifrontiers Lab at @ResearchU
Our focus is on developing self-improving agentic systems, agents that learn through interaction with humans and other agents, coordinate and collaborate, and scale into complex real-world environments, covering everything from training and evaluation to deployment.
If your expertise includes agentic AI, multi-agent reasoning, continual learning, or synthetic data and evaluation, we want to hear from you!
📝 Apply from below links 📝
https://t.co/Ni8ROs0PtE
https://t.co/Ce9oY20F0C
1/ New paper! "Wait, Wait, Wait… Why Do Reasoning Models Loop?"
Under greedy/low-temp decoding, reasoning LLMs get stuck in loops repeating themselves, wasting test-time compute and sometimes never terminating!
We study why this🔁 happens and why increasing temp is a band-aid
🚀 New Personalization Benchmark 📊
Lifelong agents must track how users evolve, not just today's preferences.
HorizonBench is exactly the kind of benchmark we need: 6-month interaction histories, preference drift, and models that still reach for who you used to be.
🧵👇
Generator, benchmark, and mental state graphs all released. Build your own splits, probe specific failure modes, all without real user data.
We hope the HorizonBench benchmark and data generator open up long-horizon personalization as a tractable research problem for the community.
📄https://t.co/JvwlPAMiqo
🤗https://t.co/ceKWxzqeZD
💻https://t.co/ATUSRLUxCP
This project couldn't have been possible without my amazing collaborators and mentors @bvp22294@Keremoktar@Diyi_Yang@tsvetshop@real_asli 🥰💙
🧠Managing context just got a lot smarter. 🦾
My colleagues at AI Frontiers at Microsoft Research just dropped Memento, and we are open-sourcing the whole stack. Give your models better memory! 💾✨
Teach your own model to manage its context with Memento. We're open sourcing everything.
Brought to you by AI Frontiers, a boutique lab within Microsoft Research
Teach your own model to manage its context with Memento. We're open sourcing everything.
Brought to you by AI Frontiers, a boutique lab within Microsoft Research
I don't think I've said this yet but 1.5 years in and Microsoft Research is an incredible place to do research. Well resourced, real chance to influence product, and a rare combo of freedom and direction. Extracting a lot of units of Dimitris/day :)
Does personalization really require endless history? 🤔
While RL is incredibly powerful, we found a beautifully simple angle: exploiting preference priors!✨
So proud of our new work co-led by brilliant @stellalisy & @avibose22!
PEP adapts to you in just 1-2 questions! 👇🧵
Personalization assumes you need history with a user. What if you don't?
Cold-start is hard: each task&user has many preference dimensions, but each user only cares about a few.
A few strategic questions is all you need, if u know how preferences correlate across population👉🏻🧵
@_Hao_Zhu@ArpandeepKhatua@SAP@michaelryan207@jiaxin_pei@Diyi_Yang Great work and such a timely benchmark. Infact, we have studied two agent collaboration in in our recent NeurIPS paper. We will try to use your benchmark . Hope you cite our work: https://t.co/ufy3wQFKIa
Love seeing Matrix getting attention! 🚀 Thank you
@omarsar0
We’re taking it further at #NeurIPS2025 next week, showing how it powers our Collaborative Reasoner agents that learn to reason together. https://t.co/mMukN7pWt2
Cool paper from Meta.
And another excellent application of multi-agent systems.
(bookmark it)
Training modern AI models requires massive amounts of high-quality data.
However, the bottleneck isn't just quantity. The data is just not diverse enough. Single models generating synthetic data tend to produce homogeneous outputs, repeating patterns, and lacking the nuanced variety found in human-created datasets.
This new research from Meta introduces Matrix, a peer-to-peer framework where multiple AI agents collaboratively generate synthetic training data through decentralized interactions.
Matrix achieves 2–15× higher data generation throughput under identical hardware resources, without compromising output quality.
TL;DR: Instead of one model producing data, specialized agents play distinct roles and interact with each other. One asks questions, another responds, a third evaluates quality. These multi-turn conversations capture complex reasoning and diverse perspectives.
What makes Matrix different: no central coordinator. Agents communicate directly in a fully decentralized architecture. This enables scalability without infrastructure bottlenecks.
The framework operates through role-based conversation protocols, multi-turn interaction patterns, and built-in quality filtering at each stage. Only data meeting quality thresholds makes it into the final training set.
Multi-agent collaboration produces more diverse synthetic data than single-model approaches. The resulting datasets improve downstream model performance across reasoning and instruction-following benchmarks.
Our team at FAIR is hiring PhD research interns for 2026 on the topics of multimodal multi-agent learning. If you are interested, feel free to DM me or directly apply using the link below!
https://t.co/JrHoDAPDnP