@coreylynch Super impressive! But I’m surprised you don’t just use Sora to generate video contextualized on the conversation and recent video input then judt translate from the video to robot movement. Maybe I’m overestimating Sora but seems it would remove the need to train on new tasks.