1/6 Diffusion models are scaling up, but deploying a massive, monolithic network uniformly across the entire generative timeline is inherently inefficient.
Introducing Complexity-Balanced Splitting (CBS): a principled framework that allocates capacity exactly where needed!👇🧵
🎉 Happy to share our latest work: Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
We generate camera-aligned, topology-consistent 4D meshes from videos x13-x200 faster than prior work and scale to videos up to 16× longer without degrading mesh quality.
Our paper:
"LaMI: Augmenting Large Language Models via Late Multi-Image Fusion"
has been selected for an Oral Presentation at #ACL2026!
LaMI boosts LLM visual commonsense by generating complementary images from a text prompt and late-fusing their evidence into the prediction
🧵
Cool paper by Tzachor et al. asking what if your Multimodal LLM already has better video embeddings than trained Video Models?
VidVec finds strong video–text reps in intermediate layers, and with text-only “in-context” optimization achieves SoTA across MSR-VTT/MSVD/VATEX/DiDeMo
OmnimatteZero. Training-free video editing. Pulls foreground objects (shadows/reflections) and reconstructs backgrounds in real-time.
- LTX-Video/Wan2.1, ~24fps on an A100.
Finally, clean background without manual rotoscoping.
https://t.co/8Oh1icqPZE
I felt like Indiana Jones unearthing a hidden treasure with this. 🤠 NVIDIA's new AI tech deletes the un-deletable - shadows on grass and more. All this in real time! Full video: https://t.co/bQIdQAuyyc
🚀 Excited to share our new paper: “Fast Autoregressive Video Diffusion & World Models with Temporal Cache Compression & Sparse Attention.”
We address attention bottlenecks in auto-regressive video diffusion, enabling ×5–×10 speedup and constant memory over long rollouts.
Excited to share our paper “When Are Concepts Erased from Diffusion Models?” at @NeurIPSConf!
We introduce two conceptual models for erasure mechanisms in diffusion models, and a suite of probes to recover supposedly forgotten concepts.
Project website: https://t.co/pKQmjEASHK
Excited to share this has now been accepted at #NeurIPS2025 as a position paper (<6% acceptance)!🎉
We advocate for systematically studying entire model populations via weight-space learning, and argue that this requires charting them in a Model Atlas.
@NeurIPSConf#NeurIPS
🧵👇
🤖 AI for finding a needle in a haystack:
🚀 We're excited to share our #NeurIPS2025 paper: "Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization".
Key contributions:
⭐ We study and analyze the problem of SoIR.
⭐ We introduce new benchmarks specifically designed for SoIR.
⭐ We propose MaO, our novel state-of-the-art method that leverages explainability maps to refine object representations into a single image descriptor.
[1/10] 🤔 What if you wanted to generate a 3D model of a “Bolognese dog” 🐕 or a “Labubu doll” 🧸?
Try it with existing text-to-3D models → they collapse.
Why? These concepts are rare or new, and the model has never seen them.
🚀 Our solution: MV-RAG
See details below ⬇️
[1/6] 🎬 New paper: Story2Board
We guide diffusion models to generate consistent, expressive storyboards--no training needed.
By mixing attention-aligned tokens across panels, we reinforce character identity without hurting layout diversity.
🌐 https://t.co/aRG81nu5qK
I'm very excited to announce our #SIGGRAPH2025 workshop:
Drawing & Sketching: Art, Psychology, and Computer Graphics 🎨🧠🫖
🔗 https://t.co/iagFGSxMrZ
📅 Sunday, August 10th
Join us to explore how people draw, how machines draw, and how the two might draw together! 🤖✍️
🎉 Our paper DOVE 🕊️ has been accepted to #ACL2025 Findings!
DOVE 🕊️ is a massive collection (250M!) of LLM outputs across different prompts, domains, and models, aimed at democratizing LLM evaluation research!
Thanks to all collaborators!
Paper:
https://t.co/BxW5ZQuoKB
Think your latent-noise diffusion watermarking method is robust? Think again!
We show that they are susceptible to adversarial attacks that only require one watermarked example and an off-the-shelf encoder. This attack can forge and remove the watermark with very high accuracy