Could this be the ViT moment for 3D scene understanding? 🚀
We revisit the good old Transformer architecture and apply it to 3D scene understanding with minimal modifications. #Volt ⚡️
Project page: https://t.co/1HYyTrtdak
Arxiv link: https://t.co/imcuvM6j0X
(1/5)
🔨 Presenting HAMMER at #CVPR2026 today!
📍 Poster Session 4 · ExHall A · Board #221
MLLMs for intention-driven 3D affordance grounding. Code, weights & data all open.
Keen on 3D scene & object understanding? Come say hi!
🔗 https://t.co/mQ3vp10ZWa
Excited to present our #CVPR2026 Highlight paper tttLRM today!
We design a LRM with Test-time Training that supports high-res, long sequence and feedforward/online 3D reconstruction with linear complexity.
📍 ExHall A #39
🕞 6/7, 3:30 PM
🎥 Video: https://t.co/eQB7ngurfk
📢 Excited to share MoScale, the first next-scale prediction framework for text-to-motion generation!
If you’re attending #CVPR, come visit our poster in Session 3, #199, 11:45 AM. We’d love to discuss ideas and feedback!
Paper and code can also be found at https://t.co/xw72no8AXu.
🔨 Presenting HAMMER at #CVPR2026 today!
📍 Poster Session 4 · ExHall A · Board #221
MLLMs for intention-driven 3D affordance grounding. Code, weights & data all open.
Keen on 3D scene & object understanding? Come say hi!
🔗 https://t.co/mQ3vp10ZWa