I'm going to present Fast-AR at #ICML26 tomorrow (Tuesday) @ 2pm - poster #1204!
Come and say hi 👋
If you're around this week, feel free to DM me.
Project page: https://t.co/M7SqMFAHfj
Details below ⬇️
1/7
New paper 🧵 ScenA generates a multi-speaker audio scene: overlapping speech, laughter, real room noise, from a text description and a few reference voices.
When trained in the obvious way, it ignores the text and decides who speaks on its own. We found out why and fixed it.
We find a shortcut in training reference-conditioned audio flow matching: at low noise levels, the model assigns speakers based on acoustic similarity rather than text. Biasing timestep sampling toward higher noise removes this shortcut and restores text-driven speaker assignment
Even today, with powerful image editing models, making fine-grained structural changes to 3D shapes remains a major challenge.
In our new #SIGGRAPH2026 paper, Prox-E, we use primitive-based abstraction to leverage VLMs for precise, reasoning-based 3D editing!
👇
🚀Excited to present our new paper that has been accepted to #WACV2026!
Text-to-image models often fail at simple spatial tasks, like placing a dog to the right of a teddy bear.
Our solution: Learn-to-Steer.
We learn a loss function directly from attention maps and apply it during inference.
This work was done together with @AtzmonYuval and @GalChechik
📰arXiv: https://t.co/flonjHFovr
🌐Project page: https://t.co/ReSmHuUzl3
📽️Video: https://t.co/MYscutOv8L
🚀 Excited to share our new paper: “Fast Autoregressive Video Diffusion & World Models with Temporal Cache Compression & Sparse Attention.”
We address attention bottlenecks in auto-regressive video diffusion, enabling ×5–×10 speedup and constant memory over long rollouts.
Model merging is a game-changer for multitask learning without retraining. But why do some models merge perfectly while others crash? In our new paper we define and investigate Mergeability.
#LLMs#MachineLearning#NLProc#ModelMerging
Thrilled to share that two papers got into #NeurIPS2025 🎉
✨ FlowMo (my first last-author paper 🤩)
✨ Revisiting LRP
I’m immensely proud of the students, who not only led great papers but also grew and developed so much throughout the process 👇
🎉 I am excited to present our new paper!
Our paper improves personalization of text-to-image models, by adding one special cleaning step on top of existing personalized models.
With just a single gradient update (~4 seconds on an NVIDIA H100 GPU) and a single image of the target concept, our method improves both text alignment and image alignment. For example, it improves LoRA by (+7% / +14%). This is achieved by adding new loss terms and taking into account the prompt and seed.
This work was done together with @dvir_samuel and @GalChechik.
🌐 Paper page: https://t.co/5KeXClcVd3
📄 arXiv paper: https://t.co/U1TXwt35EJ
More details in the comments below.
Can AI coscientists help automate ablation planning?
To test this, we created AblationBench, a benchmark suite to evaluate models on ablation planning in empirical AI research.
The results? Even the best current models recover only 29% of the original ablations on average ⬇️
So Israel is expected to surrender to Hamas & feed them even though Israeli hostages are being starved? Did UK surrender to Nazis and drop food to them? Ever heard of Dresden, PM Starmer? That wasn’t food you dropped. If you had been PM then UK would be speaking German!
In his own words, senior Hamas terrorist
Ghazi Hamad admits:
“The recognition of a Palestinian state is one of the fruits of October 7״.
Don’t let terror collect its reward.
My God, the only smiling picture of Greta Thunberg ever is from today when she saw Israelis. So, Israel makes you smile? The Israelis welcomed her with no filters and extra hummus. Israel handled it like pros. No drama, no terrorists, just aid delivered the right way without the floating selfie circus and influencer Olympics. They wanted this to be used against Israel, but like it or not, it showed humanity, law, and order.
Israel won this.
#Madleen