Autoregressive models dominate, but what if we treat multimodal generation as discrete order agnostic iterative refinement? Excited to share our systematic study on the design space of Tri-Modal Masked Diffusion Models (MDMs). We pre-trained the first Tri-Modal MDM from scratch on (text,), (image, text), and (audio, text). The same model can do ASR, TTS, T2I, captioning and native text generation.
What I'm the most proud of in this work is the scientific rigor. Over 3,500 training runs. Principled hyperparameter transfer. Honest results. Carefully controlled ablations across multiple different axis of entanglement.
A thread on our empirical findings (arXiV: https://t.co/ZJHmk5OLsl)
Our work on fine-grained control of LLMs and diffusion models via Activation Transport will be presented @iclr_conf as spotlight✨Check out our new blog post https://t.co/dAJQtcETNX
I am so excited to be attending my first @NeurIPSConf this year!! Hit me up if you would like to chat.
On Sunday, I will be at the MINT workshop on model interventions. Join us to understand the inner workings of foundation models.
https://t.co/LOKZwoEzut
Thrilled to share the latest work from our team at @Apple where we achieve interpretable and fine-grained control of LLMs and Diffusion models via Activation Transport 🔥
📄 https://t.co/TYlwxarrWx
🛠️ https://t.co/gciUcwRNqd
1/9 🧵
Excited to announce that our workshop "I Can't Believe It's Not Better: Challenges in Applied Deep Learning" has been accepted at #ICLR2025! 🎉 Stay tuned for updates on the agenda, and paper submissions.
Do you work on foundation models?
➡️Submit to the @NeurIPS2024 🍃MINT 2024: Workshop on Foundation Model Interventions
We feature a fantastic speaker lineup with @AtticusGeiger@hardmaru@JacobSteinhardt@viegasf, space for discussions, and more!
https://t.co/e83nVihujb