Interested in improving the inference quality of diffusion models? Checkout our #FreeInit poster at No.226 from 10:30am to 12:30pm tomorrow (Tuesday 1 Oct). Will be happy to discuss video diffusion models, noise initialization and more! #ECCV2024
📢Another Free Lunch for GenAI📢
We propose #FreeInit, a sampling strategy to improve temporal consistency of video generation at inference time, requiring no training and can be plugged into any diffusion model
- Project: https://t.co/6KNCrs5C8p
- Code: https://t.co/olvF6oiYs0
Interested in improving the inference quality of diffusion models? Checkout our #FreeInit poster at No.226 from 10:30am to 12:30pm tomorrow (Tuesday 1 Oct). Will be happy to discuss video diffusion models, noise initialization and more! #ECCV2024
📢Another Free Lunch for GenAI📢
We propose #FreeInit, a sampling strategy to improve temporal consistency of video generation at inference time, requiring no training and can be plugged into any diffusion model
- Project: https://t.co/6KNCrs5C8p
- Code: https://t.co/olvF6oiYs0
Carry your cool backpack 🎒 and come to visit us in the October 1 and 2 afternoon sessions!
Octopus 🐙
- a vision-language model for robots 🤖 or GTA/Minecraft game play 🎮
- We solve it via training the VLM into a good programmer with sight. It calls correct functions and executes.
FunQA 💃🪩🕺
- a testbed to see whether video language models can understand stupid 🤪funny🔥 viral videos.
- Let’s see how your VLMs understand these fun videos!
FreeInit: Bridging Initialization Gap in Video Diffusion Models
paper page: https://t.co/RhCifC8bnd
Though diffusion-based video generation has witnessed rapid progress, the inference results of existing models still exhibit unsatisfactory temporal consistency and unnatural dynamics. In this paper, we delve deep into the noise initialization of video diffusion models, and discover an implicit training-inference gap that attributes to the unsatisfactory inference quality. Our key findings are: 1) the spatial-temporal frequency distribution of the initial latent at inference is intrinsically different from that for training, and 2) the denoising process is significantly influenced by the low-frequency components of the initial noise. Motivated by these observations, we propose a concise yet effective inference sampling strategy, FreeInit, which significantly improves temporal consistency of videos generated by diffusion models. Through iteratively refining the spatial-temporal low-frequency components of the initial latent during inference, FreeInit is able to compensate the initialization gap between training and inference, thus effectively improving the subject appearance and temporal consistency of generation results. Extensive experiments demonstrate that FreeInit consistently enhances the generation results of various text-to-video generation models without additional training.
We also find this strategy has broader applications. For example in T2I, FreeInit can help SDXL generate very dark/bright images with better text alignment (see below). Further support for #VideoCrafter2 and #SDXL will be updated soon.
[2/3]
🔥Video Generation with Image Prompts🔥
#CVPR2024 We propose *video generation with image prompts* 📽️VideoBooth📽️, providing more direct content control beyond text prompts @CVPR
- Project: https://t.co/V2GxUoYm9J
- Paper: https://t.co/2LxclfWWRF
- Code: https://t.co/F3drBPIW03
FreeInit : Bridging Initialization Gap in Video Diffusion Models @Gradio demo is out on @huggingface
demo: https://t.co/4lHfjgdbNV
run with docker: https://t.co/ZdMJ3kbKqZ
duplicate space with private gpu and no queue: https://t.co/GhrFGeqxOb
The @Gradio demo for #FreeInit is now out on @huggingface🤗! Feel free to try it out and play around with the parameters :)
HF demo link: https://t.co/ePaHDjMsyA
FreeInit: Bridging Initialization Gap in Video Diffusion Models
paper page: https://t.co/RhCifC8bnd
Though diffusion-based video generation has witnessed rapid progress, the inference results of existing models still exhibit unsatisfactory temporal consistency and unnatural dynamics. In this paper, we delve deep into the noise initialization of video diffusion models, and discover an implicit training-inference gap that attributes to the unsatisfactory inference quality. Our key findings are: 1) the spatial-temporal frequency distribution of the initial latent at inference is intrinsically different from that for training, and 2) the denoising process is significantly influenced by the low-frequency components of the initial noise. Motivated by these observations, we propose a concise yet effective inference sampling strategy, FreeInit, which significantly improves temporal consistency of videos generated by diffusion models. Through iteratively refining the spatial-temporal low-frequency components of the initial latent during inference, FreeInit is able to compensate the initialization gap between training and inference, thus effectively improving the subject appearance and temporal consistency of generation results. Extensive experiments demonstrate that FreeInit consistently enhances the generation results of various text-to-video generation models without additional training.
Thanks @_akhaliq for sharing! 🎉
We propose #FreeInit to bridge the training/inference gap of video diffusion models, improving temporal consistency without finetuning.
Project: https://t.co/n8J2iq3O0d
Code: https://t.co/n2wZZ83eU1
Video: https://t.co/YWtceq2eJ9
FreeInit: Bridging Initialization Gap in Video Diffusion Models
paper page: https://t.co/RhCifC8bnd
Though diffusion-based video generation has witnessed rapid progress, the inference results of existing models still exhibit unsatisfactory temporal consistency and unnatural dynamics. In this paper, we delve deep into the noise initialization of video diffusion models, and discover an implicit training-inference gap that attributes to the unsatisfactory inference quality. Our key findings are: 1) the spatial-temporal frequency distribution of the initial latent at inference is intrinsically different from that for training, and 2) the denoising process is significantly influenced by the low-frequency components of the initial noise. Motivated by these observations, we propose a concise yet effective inference sampling strategy, FreeInit, which significantly improves temporal consistency of videos generated by diffusion models. Through iteratively refining the spatial-temporal low-frequency components of the initial latent during inference, FreeInit is able to compensate the initialization gap between training and inference, thus effectively improving the subject appearance and temporal consistency of generation results. Extensive experiments demonstrate that FreeInit consistently enhances the generation results of various text-to-video generation models without additional training.
@DigThatData@_akhaliq Yeah it looks like this partial renoising strategy may have some similar functionalities. Very interesting. Thanks for discussion! I think the training-inference gap is the core problem, and other better approaches can be developed to address this.
@DigThatData@_akhaliq In the ablation, we show that using multiple passes w/o NR can indeed improve consistency, but not as good. From my experience, multiple passes cannot fix some extreme inconsistencies, and this is exactly where the FFT operations takes part in. Please refer to Fig.A11 for details
📢Another Free Lunch for GenAI📢
We propose #FreeInit, a sampling strategy to improve temporal consistency of video generation at inference time, requiring no training and can be plugged into any diffusion model
- Project: https://t.co/6KNCrs5C8p
- Code: https://t.co/olvF6oiYs0
@Truongtoc1980@_akhaliq Also, our core observation is: the quality of the initial noise's low frequency components is crucial for consistency. FreeInit is just one concise attempt, I think more approaches can be developed to address this problem.