Can LLM agents coordinate in long-horizon, open-ended worlds?
We evaluate 13 modern LLMs in a new benchmark where agents must work together to explore, communicate, trade resources, craft tools, build structures, and fight mobs.
TL;DR: Most agents struggle, averaging only ~6% normalised return. Yet on the hardest setting, zero-shot Gemini 3.1 Pro performs comparably to the best MARL agent trained for 1 billion environment steps.
More broadly, we find coordination is a distinct bottleneck beyond long-horizon task competence, with communication having the largest effect in our harness ablations. 🧵👇
Just gave a presentation at an #ICML workshop on our Molten Pot: Evaluations and Datasets for Social Offline Reinforcement Learning! 🤖💬🤖Really happy that this paper received the Runner-up Best Paper Award 🥈
And congrats to co-lead author @hedman_marcel 🚀
FairBED: A Bayesian Experimental Design Approach to Gathering Fairer Data
Marcel Hedman, Emily Alger, Brieuc Lehmann, Chris Holmes, Tom Rainforth
https://t.co/id91bBCt0V [𝚜𝚝𝚊𝚝.𝙼𝙻 𝚌𝚜.𝙻𝙶]
We’re open-sourcing the models, weights and training code for PERSIST, our 1.7B parameter 3D world model!
By modelling a dynamic 3D world state instead of relying on pixel histories, PERSIST generates experiences that remain spatially and temporally coherent over thousands of steps.
Tom, @kaixin20578389 and I will present PERSIST at @icmlconf in Seoul next month. See you there!
Links in thread⬇️
#WorldModels #ICML2026
Register for our next ‘Fellows’ Spotlight’ seminar led by @ClaudeFormanek (@UCT_news) exploring his recent work on 'Benchmarking Offline RL in Mixed-Motive Social Settings', going live 21 May, 16:00 UTC. Link below.
Ever wondered why presenting more facts can sometimes *worsen* disagreements, even among rational people? 🤔
It turns out, Bayesian reasoning has some surprising answers - no cognitive biases needed! Let's explore this fascinating paradox quickly ☺️
We came up with a really simple way to train flow matching (diffusion) policies with offline RL! Flow Q-learning from @seohong_park uses a distillation (reflow-like) scheme to train flow matching actor, and works super well!
Check it out: https://t.co/TYYXGuyAgI
It’s been a dream of mine since I started in ML to see autonomous agents conduct research independently and discover novel ideas! 💡 Today we take a large step towards making this a reality.
We introduce *The AI Scientist*, led together with @_chris_lu_ and @RobertTLange. [1/N]
We're announcing a multi-year partnership with @MistralAI, as we build on our commitment to offer customers the best choice of open and foundation models on Azure.
🚨 New paper! 🚨
We introduce System 2 Attention (S2A).
- Soft attention in Transformers is susceptible to irrelevant/biased info
- S2A uses LLM reasoning to generate what to attend to
Improves factuality & objectivity, decreases sycophancy.
https://t.co/pY3tr0LrUW
🧵(1/5)
I’m told mine is a contrarian view on the events of the last few days, so here goes…
Contrary to what @kevinroose and others have written, Microsoft was not a winner of the events of the last few days around #OpenAI. They were in a much better place on Friday morning last week than they are today. Friday morning they had invested ~$11B in OpenAI and captured most of its upside while still having enough insulated distance to allow @BradSmi to claim things to regulators like “ChatGPT is more open than Meta’s Llama” and to allow any embarrassing LLM hallucinations or other ugliness to be OpenAI’s problem, not Microsoft’s.
Sunday morning they were at their lowest point. There was real risk they lose all their ~$11B investment and look like absolute fools for making that size of investment without any real governance controls. Very smart people who have followed the news carefully, including some big fund managers who hold $MSFT, are still pinging me asking: how was that even possible?
Today they are better off than Sunday morning, but far, far worse off than they were Friday morning. Sure they will likely hire a bunch of the OpenAI team. But that doesn’t get them much they didn’t have before, and it comes with a ton of new reputational risk (they now own responsibility for any hallucinations or other ugliness) and execution risk (see DeepMind within Google for how all the smartest people in AI can still get stymied by the bureaucracy of a giant company).
I think the chances of the senior OpenAI folks still being at Microsoft in 3 years is asymptotically approaching zero. Where the independence and clear mission of OpenAI was exactly what could have kept that group of incredible talent motivated and aligned over the long term, making Office365 spreadsheets a bit more clever isn’t something that rallies a team like their’s. Sure they’ll try and have some level of independence, but the machinery of a trillion dollar+ business software behemoth is hard to not get caught up in and ground out by.
This was a very bad weekend for Microsoft (and, for that matter, OpenAI). I don’t see any clear winners. (Maybe @elonmusk or @benioff — indirectly?) It could have been worse, but it has still been a disaster for everyone directly involved. We’re not at the end of this story. But don’t see a lot of ways in the short term it gets better for Microsoft. And really hard to see how it could get better even over the long term than it looked for them and OpenAI Friday morning last week.
A Google employee has questioned the sentience of a large language model he developed
The claim seems very unfounded but raises interesting questions over what we would need to see for us to believe it were sentient...
https://t.co/8CjiNdvjOE
#sentientAI