Do we need separate models for image editing and conditional generation?
We introduce🌟EditAR, a unified autoregressive model for diverse tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image.
https://t.co/9Zsq47laUH
More details🧵👇
I will join Tsinghua University, College of AI, as an Assistant Professor in the coming month. I am actively looking for 2026 spring interns and future PhDs (ping me if you are in #NeurIPS).
It has been an incredible journey of 10 years since I attended an activity organized by Tsinghua University and decided to change my undergraduate major from Economics to Computer Science, inspired by one of the teammates. During the 10 years, I met with appreciation of many wonderful researchers/professors who led me to continued growth. 🐿️
My research focus will continue to be AI & Robotics, with a specific emphasis on Interactive Embodied Intelligence. You can check my homepage to learn more: https://t.co/6hHsc62x0A.
I am currently local to San Diego and will be attending #NeurIPS. Please ping me over WeChat or Email if any old or new friends are interested in having a coffee chat! (Really looking forward to meeting as many friends as possible at #NeurIPS)
[The photo is one of the places that I will miss a lot in the US]
How to generate billion-scale manipulation demonstrations easily? Let us leverage generative models! 🤖✨
We introduce Dex1B, a framework that generates 1 BILLION diverse dexterous hand demonstrations for both grasping 🖐️and articulation 💻 tasks using a simple C-VAE model.
It's challenging, but so rewarding! Thank you, @xiaolonw, 🥰 for being a steady source of support and mentorship. I am especially grateful for the freedom you gave me to follow my curiosity. I also feel lucky to have shared this journey with such an inspiring group of labmates!
Congratulations to the graduation of @Jerry_XU_Jiarui@JitengMu@RchalYang@YinboChen !
I am excited for their future journeys in industry:
Jiarui -> OpenAI
Jiteng -> Adobe
Ruihan -> Amazon
Yinbo -> OpenAI
🥳 EditAR code is released! Welcome to check it out.
👉Presenting EditAR at #CVPR2025!
(Friday afternoon, Jun 13, 4:00pm-6:00pm, Hall D #242)
Code: https://t.co/R2yxJEQst5
Project: https://t.co/mkbMUQdph3
Do we need separate models for image editing and conditional generation?
We introduce🌟EditAR, a unified autoregressive model for diverse tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image.
https://t.co/9Zsq47laUH
More details🧵👇
Test-Time Training (TTT) is now on Video! And not just a 5-second video. We can generate a full 1-min video!
TTT module is an RNN module that provides an explicit and efficient memory mechanism. It models the hidden state of an RNN with a machine learning model, which is updated via gradient descent. Combined with a Diffusion Transformer, we are able to generate a 1-min Tom and Jerry cartoon.
Enjoy our video with input script (not seen before):
Jerry happily eats cheese in a tidy kitchen until Tom playfully takes it away, teasing him. Annoyed, Jerry packs his belongings and leaves home, dragging a small suitcase behind him. Later, Tom notices Jerry's absence, feels sad, and follows Jerry's tiny footprints all the way to San Francisco. Jerry sits disheartened in an alleyway, where Tom finds him, gently offering cheese as an apology. Jerry forgives Tom, accepts the cheese, and the two return home together, their friendship restored.
Today's visual generative models are mere stochastic parrots of imagery, much like early language models, which could only statistically mimic short sentences with little reasoning. In contrast, modern large language models (LLMs) can comprehend long documents, keep track of dense information contexts, reason about complex problems, all the while providing natural human interactions.
Reve’s mission is to invent the future of intent-driven visual creation. Capturing creative intent requires advanced machine understanding of natural language and other interactions. Turning this intent into compelling visuals calls for interactive systems that have a deep understanding of the visual world they generate, so they can iteratively amend it.
Our vision is to build a new semantic intermediate representation that both a human and a machine can understand, reason about, and operate on, together with a probabilistic render to decode that representation.
Today, we are thrilled to share our first renderer, Reve Image, with the world!
(also, it’s [ʀɛv], from “rêve” 🙂)
Excited to come out of stealth at @reveimage!
Today's text-to-image/video models, in contrast to LLMs, lack logic. Images seem plausible initially but fall apart under scrutiny: painting techniques don't match, props don't carry meaning, and compositions lack intention. (1/4)
Meet our first general-purpose robot at @DexmateAI
https://t.co/RzWzxjP3Xu
Adjustable height from 0.66m to 2.2m: compact enough for an SUV, tall enough to reach those impossible high shelves. Powerful dual arms (15lbs payload each) and omni-directional mobility for ultimate versatility.
More videos showcasing Vega in action coming this week! Stay tuned to see this robot partner.
Do we need separate models for image editing and conditional generation?
We introduce🌟EditAR, a unified autoregressive model for diverse tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image.
https://t.co/9Zsq47laUH
More details🧵👇
🐅 Want to rig your favorite meme character?
Try “RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets”!
✨RigAnything is a transformer-based model that sequentially generates skeletons without predefined templates. It creates high-quality skeletons for shapes in any pose, completing rigging in under 2 seconds per shape. 🧵(1/n)
Introducing “Diffusion Autoencoders are Scalable Image Tokenizers” (DiTo).
We show that with proper designs and scaling up, diffusion autoencoders (a single L2 loss) can outperform the GAN-LPIPS tokenizers (hybrid losses) used in current SOTA generative models. (1/4)
Autoregressive models are picking up for image generation, but how about image editing?
Given an image and a language description, we train one Autoregressive Transformer to do ANY EDITING.
Do we need separate models for image editing and conditional generation?
We introduce🌟EditAR, a unified autoregressive model for diverse tasks, e.g., image editing, depth-to-image, edge-to-image, segmentation-to-image.
https://t.co/9Zsq47laUH
More details🧵👇
EditAR produces photo-realistic results and offers substantial sample diversity. Our experiments show that it achieves even better FID 🌟compared to various specialized diffusion methods.