Excited to share RLE-Bench, a benchmark we’ve been building over the past few months. We envision agents that can perceive, control robots, develop policies, and design their own hardware. RLE-Bench measures how close today’s coding agents are to that vision.
GPT-6 Astra is impressive, but agentic robotics ≠ just control.
Physical agents should control robots, learn new skills, and design their own hardware.
So we introduce RLE-Bench: 48 everyday robotics engineering tasks spanning closed-loop control, policy learning, perception, and mechanical design.