Lucida from ByteDance
Cool project! It transforms indoor video into editable 3D scenes; reconstructs objects as simulation-ready assets
- VLM-based object detection and referring descriptions.
- good at pose alignment
- high-fidelity real-to-sim digital twins.
- multi-view masks, boxes, partial point clouds
- Seed3D 2.0 for image-to-3D asset gen
https://t.co/uUlcInA2XL
SIMART: Decomposing monolithic meshes into sim-ready articulated assets
A unified MLLM framework using Sparse 3D VQ-VAE reduces tokens by 70% to enable part-level decomposition and kinematic prediction for physics-based robotic simulation.
Thrilled to share that our paper SIMART has been accepted to #SIGGRAPH2026! 🚀
Check out the full demo and details:
Project Page: https://t.co/cQ47aHYVU9
ArXiv: https://t.co/IiFfCWdcQh
Demo Video: https://t.co/fHED9yF1uL
#EmbodiedAI#MLLM#SIGGRAPH2026#AI
Thanks to @_akhaliq for sharing our work on #MeshCoder !
Project Page: https://t.co/aJfFQEy3mb
Paper: https://t.co/THLCu6PZmk
Code: https://t.co/xyK6g0byua
GUAVA: Generalizable Upper Body 3D Gaussian Avatar
Contributions:
• We propose GUAVA, the first framework for generalizable upper-body 3D Gaussian avatar reconstruction from a single image. Using projection sampling and inverse texture mapping, GUAVA enables fast feed-forward inference to reconstruct Ubody Gaussians from the image.
• We introduce an expressive human template model with a corresponding upper-body tracking framework, providing an accurate prior for reconstruction.
• Extensive experiments show that GUAVA outperforms existing methods in rendering quality and significantly outperforms 2D diffusion-based methods in speed, offering fast reconstruction and real-time animation.
4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models
Contributions:
• We introduce 4D LangSplat for open-vocabulary 4D spatial-temporal queries. To the best of our knowledge, we are the first to construct 4D language fields with object textual captions generated by MLLMs.
• To model smooth transitions across states for objects in 4D scenes, we propose a status deformable network to capture continuous temporal changes.
• Experiential results show that our method attains state-of-the-art performance for both time-agnostic and time-sensitive open-vocabulary queries.
"LangSplat: 3D Language Gaussian Splatting." Everything including code released but the paper is not out yet 🎅
https://t.co/ohnIuZKFif
https://t.co/Ssakr5sL41
Thanks to @_akhaliq for sharing our #CVPR2025 work on #4DLangSplat!
Project page: https://t.co/96WJ7nXJXg
paper: https://t.co/gCwtiWLlJJ
Code: https://t.co/KrTr4Rg4Bc
LangSplat: 3D Language Gaussian Splatting
paper page: https://t.co/RtOpGlO8lS
ground CLIP features into a set of 3D language Gaussians, which attains precise 3D language fields while being 199 × faster than LERF
Tsinghua and @Harvardresearchers introduced #LangSplat, an advanced #AI method for 3D language fields. It uses 3D Gaussian Splatting to enhance user communication with #3Denvironments.
For more: https://t.co/HY591hSs0M