You never know how much you don’t know. Robotics RL is actually difficult. I know the high level reasonings behind the pipeline and the ML aspect but pattern matching is not enough. Why did my run improve by 2% and then crash to 100% failure? The KL term is too low for the learning rate. The critic’s gradient is too high. The physics engine isn’t accurate and causes grasping to fail. The tactile sensors are creating nans. All of the sensors are making training very slow. Does the sim actually transfer to real. The hz in sim is high compared to the inputs we get in real. The policy isn’t moving the robot. Too little generalization. The entropy was turned to 0 because it was most efficient to do nothing because of a simple reward error and we kept it for 3 days. Physics sim and had a fp32 vs fp64 replay error. Don’t even get me started with the numerous lies Astra told me that cost me days and also the thousand I spent in credits. I’m not even talking about reward shaping, the real aspect, or how painful it was to do data collection. That’s exactly why Wendy decided that we need to standardize development and make it really easy for everyone to deploy and test out policies. Most companies and people I talked to are not the goats at all subjects (except Wendy people ;)). Most people have their niche areas and get really good at it. The people that are RL experts and know hardware very well and specifically robotics are incredibly rare. But thankfully, it’s the most fun field in the world. Happy building everyone ;)