Feed-forward 3D reconstruction methods typically predict pointmaps in camera-centric frames. But why should a camera's arbitrary orientation define the coordinate system?
We introduce G3T, a transformer that predicts pointmaps in gravity-aligned frames. Regardless of input image orientation, our method always produces upright pointmaps (see demo).
We leverage this uprightness to create G3T-Long, a submap-based reconstruction method that improves robustness on long-sequence 3D reconstruction (more on that below).
Interactive demos, code, and model weights are available on our project page.
Excited to share that EasyV2V has been accepted to CVPR 2026! Data is the fuel. EasyV2V introduces a comprehensive and scalable video editing data pipeline. Curious how different types of data impact video editing performance? Check out our project page:
https://t.co/IsQkUoiVE2
🚀New paper Omni-Attribute is released!
This work can isolate a specific attribute from any image and merge those selected attributes from multiple images into a coherent generation.
It’s a pleasure working with @tsaishien_chen, the amazing team at @Snap, and @junyanz89 from CMU!
I’m excited to be at NeurIPS 2025, presenting our latest work “Preventing Shortcuts in Adapter Training via Providing the Shortcuts”
📅 Wed, Dec 3, 2025
⏰ 11:00 AM – 2:00 PM PST
📍 Exhibit Hall C, D, E — Poster #4406
📄 arXiv: https://t.co/At21XrLGmb
Happy to chat about image/video personalisation, on-device diffusion models, and life beyond work!
Looking fwd to connecting with old friends and new at #CVPR2025
Heading to @CVPR 2025 in Nashville this week? So are we!
We’re proud to have 12 papers accepted — including SnapGen and 4Real-Video, both highlighted among the top 3% of submissions.
Come find us to learn more about the cutting edge work we’re doing in AI and computer vision.
📍 See you in Nashville!
Learn more: https://t.co/JQIghmmgo6
📢I am attending #CVPR2025 (Jun 11 - 14). Come to our https://t.co/vAx6y3opa9 poster to know more about how we achieved the highest ID preservation in personalization and further enables expression following in our follow ups. See you at Fri 4 - 6 pm, ExHall D Poster #326.
Can we predict 3D human posture and body pressure from a single depth and pressure image?
Check out our #CVPR2024 work, BodyMAP: Jointly Predicting Body Mesh & 3D Applied Pressure Map for People in Bed 🧵
📅 Poster: Wed 19 Jun, 10:30 a.m. Arch 4A-E #222 https://t.co/PJn29fvRi7