Attending #ECCV2026? Check out our cool work on attention sparsity in video diffusion models, CalibAtt.
We achieve up to 1.58x speedup with no training or finetuning!
https://t.co/oK4qpuYwPs
[1/6]
After 2 years, I've wrapped up a fantastic chapter at Apple. I'm incredibly thankful to my former team, I know our paths will cross again!
I've moved to the Bay Area to join Cosmos lab at NVIDIA, working on GenAI & world models. Excited for this new adventure!
TL;DR🍎🌍 ➡️ 🟩🌎
Happy to share 🌐 SP³: Spherical Priors for Plug-and-Play Restoration, where we use Spherical Encoders as fast generative priors for image restoration.
TL;DR — strong zero-shot restoration, 3-630× faster than comparable diffusion/flow methods, and useful results from the first iteration.
Language made such a big impact on visual generative models because it balances three things at once: it's intuitive to use, expressive enough to describe complex scenes, and compact enough to condition on efficiently and at scale.
But that doesn't make natural language the perfect specification tool for images. It's vague, spatially meaningless, and incomplete — some visual specifics can only be approximated in words, never fully specified.
While going straight from text to visuals was a giant leap, this paradigm will not hold. Perhaps, it already doesn't. Today, we introduce Reve 2.0, a model with a dedicated intermediate representation — layout — that sits between user intent and finished image. Layout is structured and hierarchical: it explicitly specifies objects, their positions, sizes, colors, spatial relationships and more. You can read it, understand it, edit it, control it — and so can agents.
Layout is code for image.
Introducing 🦤 DODO: Discrete OCR Diffusion Models.
This work is the result of my summer internship at Amazon and is the first to study masked diffusion models for document parsing.
OCR is special: the image already contains the answer. So why decode one token at a time?
To folks working on image restoration
Please please please read https://t.co/4W5O1DZ8ea
and report BOTH the distortion metrics AND perceptual metrics when doing evaluation.
LPIPS does NOT count as a perception metric.
Accelerate your transformer model with the new Block-Sparse-Flash-Attention! https://t.co/d8MLq91RPW
This training-free, drop-in replacement extends FlashAttention-2 with minimal code changes (CUDA Kernels Included). Paper: https://t.co/1CLMkCUWNT
📢 Another #NeurIPS, another diffusion circle!
Join us to talk about diffusion models on Friday Dec 5 at 3:30PM in San Diego! Bayside terrace outside room 11 (upstairs) ☀️🚢🌊
Please help spread the word, tell your friends! No slides, no talks, we just sit down and chat 🗣️
I’ll be at NeurIPS from Dec 4-7, DM me if you want to chat! I’m presenting four papers this year:
📍 Thursday afternoon Poster #1002
@christy_li_ will present our self-reflective interpretability agent
https://t.co/oXRkJRsxTC
⬇️⬇️
I am recruiting exceptional PhD students and Postdocs to join my research lab at the Technion. We study Deep Learning by pursuing creative, unconventional, elegant and mathematically rigorous ideas. 👇
Lots of pixel space diffusion papers are doing the rounds this week, many of them seemingly rediscovering the benefits of the multiscale structure of UNets in the Transformer era🤭
Personally, I think we'll be doing latent diffusion for a while longer. The computational efficiency gains are large enough for people to put up with it, despite the additional complexity and relative inelegance of it all.
Eventually, they will not be, and at that point people will probably switch back to pixel space diffusion to simplify their setups, and happily eat the associated efficiency cost.
When? I'm not sure! Gemini 3 Pro says 2028-2030, so we have a few more years to go. My gut feeling is that might be an optimistic estimate. Any other guesses?
I’ll be at NeurIPS 25 in San Diego to present InvFusion (https://t.co/2MjlQo3uRL), and follow up with a visit to the Bay Area to present at @Stanford and @Berkeley.
Reach out if you want to talk about new research, diffusion models, efficient attention, or just grab a coffee :)
A key challenge for interpretability agents is knowing when they’ve understood enough to stop experimenting.
Our @NeurIPSConf paper introduces a self-reflective agent that measures the reliability of its own explanations and stops once its understanding of models has converged.
🚀 Excited to share our latest research: “SILO: Solving Inverse Problems with Latent Operators”!
A surprisingly simple approach to image restoration with latent diffusion models that achieves SOTA results while being 2.5x–10x faster than prior methods.
🧵[1/7]