🎉 I am excited to present our new paper!
Our paper improves personalization of text-to-image models, by adding one special cleaning step on top of existing personalized models.
With just a single gradient update (~4 seconds on an NVIDIA H100 GPU) and a single image of the target concept, our method improves both text alignment and image alignment. For example, it improves LoRA by (+7% / +14%). This is achieved by adding new loss terms and taking into account the prompt and seed.
This work was done together with @dvir_samuel and @GalChechik.
🌐 Paper page: https://t.co/5KeXClcVd3
📄 arXiv paper: https://t.co/U1TXwt35EJ
More details in the comments below.
4/4
It is compatible with a wide range of personalization techniques (e.g., DreamBooth, LoRA, Textual Inversion) and supports various diffusion backbones, including UNet-based models (e.g., SDXL, SD) and transformer-based models (e.g., FLUX, SD3).
🎉 I am excited to present our new paper!
Our paper improves personalization of text-to-image models, by adding one special cleaning step on top of existing personalized models.
With just a single gradient update (~4 seconds on an NVIDIA H100 GPU) and a single image of the target concept, our method improves both text alignment and image alignment. For example, it improves LoRA by (+7% / +14%). This is achieved by adding new loss terms and taking into account the prompt and seed.
This work was done together with @dvir_samuel and @GalChechik.
🌐 Paper page: https://t.co/5KeXClcVd3
📄 arXiv paper: https://t.co/U1TXwt35EJ
More details in the comments below.
2/4
It achieves new state-of-the-art results by enhancing per-query generation of personalization checkpoints, with an average gain of +8% image alignment and +23% text alignment.
1/4
Our work adds a single personalization step on top of pre-trained text-to-image personalization checkpoints that is (1) specific to the prompt and noise seed, and (2) uses two loss terms based on the self- and cross-attention, capturing the identity of the personalized concept.
I am excited to present 3to4D! A method for animating user-provided 3D objects by conditioning on textual prompts to guide 4D generation. This work was done together with @Orimalca, @dvirsam, and @GalChechik .
Project page: https://t.co/Us1yk05HiY