🚀 CFP: 1st Workshop on #PhysWorldAI at #NeurIPS2026!
Explore physical geometry, physical properties, multimodal sensing & embodied AI.
🎤 Keynotes: Anandkumar, Snavely, Finn & Freeman.
📍Atlanta, Dec 12/13 | 📅 Due Sept 9
https://t.co/A0iuxtCUyc
🚀 Stream3D is accepted to CVPR-W 2026!
The first framework for long-horizon 3D asset generation from monocular streams 🌊
— turning frozen 3D generators into streaming generators via evidential memory, no retraining
📈 Better consistency across long streams
💾 Constant memory footprint
📃 https://t.co/acp9Gydt2t
@MIT@Harvard@medialab@hkust
#AI #ComputerVision #3D
Congrats to VGGT-Omega and D4TR! 🎉
Great to see 4D reconstruction gaining more attention.
VGGT-Omage sets a strong benchmark.
Page-4D (ICLR 2026) uses the same inference backbone as VGGT, with open-source code, and still performs reasonably well in indoor scenarios. ✨
https://t.co/xY1i0IQXVA
Excited to share our new roadmap:
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
What does it mean for image generation models to become truly intelligent?
HF Daily Paper: https://t.co/jnC75XxbGf
WebPage: https://t.co/cB510BFxfp
Got a paper on generative models accepted at CVPR 2026?
Share it with us at the 4th Workshop on Generative Models for Computer Vision!
https://t.co/zwzwRvD4o8
You can simply submit your accepted CVPR paper, no need to reformat!
Deadline: April 30 (AoE)
🚀 PAGE-4D is out! (ICLR 2026)
We bring VGGT into the dynamic world 🌍
— disentangling pose & geometry for 4D perception
📈 SOTA across dynamic scene benchmarks
💻 Training + evaluation code fully released
📃 https://t.co/rUSzgLJqsW
@MIT@Harvard@medialab
#AI #ComputerVision #4D
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
Abstract (excerpt):
This survey offers a comprehensive review of feed-forward techniques for u3D reconstruction and view synthesis. It provides a taxonomy based on underlying representation architectures, including point cloud, 3D Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), etc.
We examine key tasks such as pose-free reconstruction, dynamic 3D reconstruction, and 3D-aware image and video synthesis, highlighting their applications in digital humans, SLAM, robotics, and beyond. Additionally, we review commonly used datasets with detailed statistics and evaluation protocols for various downstream tasks.
We conclude by discussing open research challenges and promising directions for future work, emphasizing the potential of feed-forward approaches to advance the state of the art in 3D vision.
Submit your extended abstract to our workshop on "Generative Models for Computer Vision"
#CVPR2025@CVPR
Authors with accepted CVPR papers are welcome to present their poster as well!
Deadline: April 25th
We also have an incredible speaker line-up!
Visual Acoustic Fields
Contributions:
• The impact locations of collected visual-sound pairs can be localized in 3D space by learning a radiance field with synchronized camera poses.
• Predicted impact sounds in our Visual Acoustic Fields accurately align with the corresponding impact locations.
• Impact regions or objects can be precisely retrieved with sound using our Visual Acoustic Fields.
@agigiraffe@jon_barron Sure, all components are modeled by Blender with some simple rotation (for objects) and particle path (for 3D Gaussians) animations.
📢Our latest preprint shows that learning global neuron shapes can help to automatically proofread connectomes and predict neuron types. https://t.co/v39Zm5Igo1 Work done in collaboration with @HHMIJanelia, @srinituraga & @HarvardVCG#connectomics#AI 🧵(1/n)