Embodied Recursive Self-Improvement (RSI) needs more than a model—it needs an environment that learns where the agent fails.
🛵 We introduce DeliveryGym—an adaptive RL environment built in Unreal Engine 5, where embodied agents learn to navigate Paris, deliver food, and earn money.
The key idea is simple: earnings are the reward, and failures shape the curriculum. DeliveryGym automatically finds where the agent struggles and generates harder tasks targeting those weaknesses.
📈 With this adaptive RL loop, Qwen3-VL-4B improves net income by 54.3% in a shift.
✨ A step toward embodied agents that fail, adapt, and improve through an environment that evolves with them.