My research focuses on offline reinforcement learning and generative models. I am currently transitioning to VLA to explore the mysteries of embodied intellige.
@yeonsumia I ran the original code with OGBench-1M. TRL/DCRL seem to use oraclerep and work better with larger datasets (e.g., 1B). Could data coverage and goal representation be bottlenecks for divide-and-conquer methods? What do you think?
@yeonsumia So DCRL uses $(x, y)$ as the goal representation in navigation tasks, right? I'm rerunning the code in the full-observation setting, but I've found that DCRL doesn't seem to achieve very high performance. Could you help me figure out what might be going wrong?