Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works!
Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training environments (using the Gym step()/reset() API), and a Reasoning Agent that learns to solve them. Resurrecting ideas from our work on UED, the Designer is trained to maximize a proxy for the Agent’s regret, computed using privileged hints.
Excited to share our latest research and open-source codebase: AgentElect! 🚀
We investigate how AI systems can leverage elections to navigate resource-based social dilemmas and drive multi-agent cooperation in common pool resource problems.
📄 Read the paper: https://t.co/4uV0H4FVJ9💻
Explore the code: https://t.co/8USGpc62Yv
#MultiAgentSystems #AI #MachineLearning #AIResearch #OpenSource