[1/6] π¬ New paper: Story2Board
We guide diffusion models to generate consistent, expressive storyboards--no training needed.
By mixing attention-aligned tokens across panels, we reinforce character identity without hurting layout diversity.
π https://t.co/aRG81nu5qK
1/7
New paper π§΅ ScenA generates a multi-speaker audio scene: overlapping speech, laughter, real room noise, from a text description and a few reference voices.
When trained in the obvious way, it ignores the text and decides who speaks on its own. We found out why and fixed it.
1/6 Diffusion models are scaling up, but deploying a massive, monolithic network uniformly across the entire generative timeline is inherently inefficient.
Introducing Complexity-Balanced Splitting (CBS): a principled framework that allocates capacity exactly where needed!ππ§΅
[6/6] Huge thanks to my collaborators for making this happen. Excited to see how the community builds on training-free consistency. More examples + code: π https://t.co/2Mawweuifj thanks to my collaborators @MatanLvy, @OmriAvr, @dvir_samuel, and @DaniLischinski
[1/6] π¬ New paper: Story2Board
We guide diffusion models to generate consistent, expressive storyboards--no training needed.
By mixing attention-aligned tokens across panels, we reinforce character identity without hurting layout diversity.
π https://t.co/aRG81nu5qK
[5/6] οΏ½οΏ½οΏ½οΏ½ To evaluate properly, we introduce the Rich Storyboard Benchmark
100 open-ended fantasy stories. 700 total panels.
Each one with detailed character actions and evolving backgrounds, designed to push generative models beyond static templates.
[4/6] π₯ How do we measure expressive storytelling?
We propose Scene Diversity, a new metric that captures variation in pose, scale, and layout across panels.
It rewards models that treat characters like real cinematic subjects, not static cutouts.