🚀 The 4-bit era has arrived! Meet #SVDQuant, our new W4A4 quantization paradigm for diffusion models. Now, 12B FLUX can run on a 16GB 4090 laptop without offloading—with 3x speedups over W4A16 models (like NF4) while maintaining top-tier image quality. #AI#Quantization. 1/7
…Sorry, we were on mute. Today we’re launching Sound Effects on Pika.
Now you can seamlessly generate and integrate sound into your videos. Either prompt the sound you want, or let Pika automatically generate it based on the content of your video.
If that sounds great, it’s because it is.
Can we construct large generative systems by composing smaller generative models? 💻 Check out this amazing paper authored by Pika Researcher @du_yilun ! 👏
https://t.co/3XnVUGI9ew
We wrote a position paper arguing that we should construct large generative models compositionally from smaller ones!
We argue that (1) it enables data/computation efficient learning (2) it enables provable generalization to unseen test distributions.
https://t.co/43r0c2xgYH
Both the article and code are now open source.
Paper: https://t.co/LjLEVqZQUM
Code: https://t.co/4518XWyhib
Authored by @LingYang_PKU, Zhaochen Yu, @chenlin_meng, @MinkaiX, @StefanoErmon, Bin Cui
Our framework can extend text-to-image generation with more conditions. Compared to ControlNet, RPG achieves significant improvements in prompt understanding and compositional semantic alignment.
Excited to announce RPG-DiffusionMaster, a joint work with Peking University and Stanford University. RPG harnesses multi-modal LLMs to master diffusion models in complex and compositional text-to-image generation/editing, achieving state-of-the-art performance.
We're hosting a private cocktail party next Thursday evening 12/14 @ NeurIPS. Enjoy an evening of drinks, food, and networking with the Pika team and AI researchers.✨
Spots limited. Apply to join at: https://t.co/YKdKlYlnPj
Pika researchers @linqi_zhou@andy_shin@chenlin_meng@StefanoErmon developed methods that achieved 4.7x speedup on text-to-3D generation
Paper: https://t.co/QsD8MFFQdR
Code: https://t.co/bQgT5MxrGX
Website: https://t.co/2MoAfyAW6d
📢Text-to-3D generation via score distillation (DreamFusion, ProlificDreamer, etc.) produces high-quality 3D assets, but can take up to 10 hours to run. We present an acceleration method for all existing approaches based on score distillation and achieves up to 4.7x speedup.