Long-horizon values depend on shorter-horizon values.
But what if those shorter-horizon values are still wrong?
We introduce DCRL: a recursive divide-and-conquer approach that learns values from short to long, substantially mitigating value-error accumulation over long horizons compared with standard GCRL methods. 🧵
Long-horizon values depend on shorter-horizon values.
But what if those shorter-horizon values are still wrong?
We introduce DCRL: a recursive divide-and-conquer approach that learns values from short to long, substantially mitigating value-error accumulation over long horizons compared with standard GCRL methods. 🧵
Long-horizon values depend on shorter-horizon values.
But what if those shorter-horizon values are still wrong?
We introduce DCRL: a recursive divide-and-conquer approach that learns values from short to long, substantially mitigating value-error accumulation over long horizons compared with standard GCRL methods. 🧵
@baymax3009 Thank you for the insightful feedback! In Figures 7 and 8, we compare our approach with a recent "explicit" quasimetric method, TMD (Myers et al., 2025).
Thanks for your interest! I believe extending DCRL to a general off-policy setting would be an important direction. Our core idea is to learn behavior values recursively while simultaneously recovering optimal values. To this end, we may need to make some changes to standard off-policy algorithms. I’m currently exploring this direction, and I hope I’ll be able to share more soon!
TL;DR: Long-horizon goal-conditioned value learning suffers from error accumulation. Recursively learning short-segment values before long-segment values mitigates this accumulation.
For more details, check out our paper and project website.
Paper: https://t.co/dyZD8d5Oad
Website: https://t.co/ejrYFoC092
w/ @YoungwoonLee
Across diverse goal-reaching tasks, DCRL substantially outperforms prior offline GCRL methods. On the five most challenging long-horizon OGBench tasks, it raises the best prior average score from 55 → 64, outperforming both flat and hierarchical baselines.
Do VLMs actually understand 3D space 🌎?
Or are they exploiting shortcuts hidden in natural images?
🚀 Excited to share our new work:
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
@NVIDIAAI × @SeoulNatlUni × @OhioStateCSE
🧵👇
🎉 Excited to share that our paper “Convergent Functions, Divergent Forms” will be presented at NeurIPS 2025🤖 in San Diego!
We present LOKI, a compute-efficient framework for co-evolving robot morphologies🦾 and control policies⚙️. LOKI discovers diverse, high-performing robot designs through shared control policies, addressing key challenges of bi-level evolution such as premature convergence and behavior collapse. It’s also significantly more efficient than Quality-Diversity methods, achieving greater diversity in locomotion behaviors and stronger transfer to various downstream tasks.
👉Project page: https://t.co/JKpy1d1hMP