Very excited to share LoGo, my internship project at @theworldlabs on post-training world models! We found that reward design matters a great deal in post-training long-horizon video gen for 3D consistency, echoing what we see in other domains, e.g. LLM reasoning. Check out more 👉 https://t.co/tLW6CVGaAk
What internal structure do experts in multimodal MoEs develop during pre-training? Can we put this structure to practical use to make model adaptation significantly faster?
In our new blog post and paper with @RaphiKang and @georgiagkioxari we discuss this and more 👇
I'll be presenting Kyvo tomorrow at @CVPR with @vanshtibrewal and @georgiagkioxari!
📍 Poster Session 3
🕥 11:45 AM to 01:45 PM (MDT)
Come by to chat. Excited to share the work and hear your thoughts!
#CVPR2026
Introducing Kyvo! 🚀 – a decoder-only LLM that aligns text, images & structured 3D scenes token-by-token.
From a single image, it reconstructs individual 3D shapes and their locations, renders & edits scenes, answers spatial questions, and more.
💻: https://t.co/WfTAgjUVTp