T2I systems look complex only because we keep adding components. How about removing them?
Introducing our latest work MiniT2I 🪄: purely flow matching objective on pixels, no auxiliary losses, all open-source data and academia-affordable compute. Everything is just designed so simple and minimal. ✨
How does a model understand and explore a world it has never seen?
We introduce 🔭VISTA🔭, a visual harness that gives a VLM long-horizon vision for reasoning in an interactive world. With Claude Opus 5.0, it reaches 100% RHAE on @arcprize's ARC-AGI-3, perfectly solving all 25 public games.
Blog post: https://t.co/bYUWwi0KoQ 🧵
Just open-sourced VBench-JAX ✨
Parity-verified VBench & VBench++ evaluation in JAX, with resumable multi-host TPU runs.
Hope it’s useful to the TPU/JAX community! I’d love to hear what works, what breaks, and what you need next 🙏
https://t.co/uMEPAtS7qD
🧚Progressive Distillation of ELF🧚
We pushed ELF, a continuous diffusion language model, even further toward few-step generation ⚡️
ELF already generates 1,024-token sequences in just 8–32 steps without distillation.
Now, with progressive distillation, ELF+PD scales all the way down to few-step and even one-step generation!
Blog: https://t.co/AX6nfu0VAK 🧵
You can always trust Kaiming's quality bar.
Writing, code, data, recipe, ckpt...
https://t.co/ogFdILwql5
https://t.co/5oUjNXtcDQ
It's a great pleasure working with all undergrads!
They're literally "postdocs" in the group.
@kevinxbwang2007@Hope7Happiness@Lyy_iiis
带带弟弟🥺
T2I systems look complex only because we keep adding components. How about removing them?
Introducing our latest work MiniT2I 🪄: purely flow matching objective on pixels, no auxiliary losses, all open-source data and academia-affordable compute. Everything is just designed so simple and minimal. ✨
Language is discrete. Language models don’t have to be.
🧚Introducing ELF🧚♀️: Embedded Language Flows—a class of diffusion models in continuous embedding space based on continuous-time Flow Matching 🧵