[1/4]🚀 Current Vision Foundation Models (VFMs), such as DINOv2, lack 3D awareness.
Our #ICLR2026 paper "Splat and Distil" addresses this by using a distillation framework with a feed-forward 3D reconstruction module, achieving SOTA performance on 3D-aware downstream tasks.
Excited to share this has now been accepted at #NeurIPS2025 as a position paper (<6% acceptance)!🎉
We advocate for systematically studying entire model populations via weight-space learning, and argue that this requires charting them in a Model Atlas.
@NeurIPSConf#NeurIPS
🧵👇
Dayani et al., "MV-RAG: Retrieval Augmented Multiview Diffusion"
Retrieval augmentation for 3D recon with multi-view diffusion models. Trains both with 3D and 2D assets during training, and uses 2D natural images for inference. Makes sense, I guess? Many things look alike!
[10/10]
Check out the full paper here 📄: https://t.co/mKWjJeo5LX
Project page: https://t.co/0oWVRrfhM4
Huge thanks to my amazing collaborators @omerbenishu & @BenaimSagie 🙌
Excited to see what the community builds on MV-RAG!
[1/10] 🤔 What if you wanted to generate a 3D model of a “Bolognese dog” 🐕 or a “Labubu doll” 🧸?
Try it with existing text-to-3D models → they collapse.
Why? These concepts are rare or new, and the model has never seen them.
🚀 Our solution: MV-RAG
See details below ⬇️
[1/6] 🎬 New paper: Story2Board
We guide diffusion models to generate consistent, expressive storyboards--no training needed.
By mixing attention-aligned tokens across panels, we reinforce character identity without hurting layout diversity.
🌐 https://t.co/aRG81nu5qK
[1/10]🚨 Introducing RewardSDS! 🚨
Standard SDS-based text-to-3D methods struggle with fine-grained alignment to user intent, often leading to artifacts or misaligned generations.
Our solution 👇