I’ll be attending #CVPR2026 and will be at the Google booth on the morning of June 5. If you’d like to chat about research or anything interesting, feel free to set up a coffee chat☕️ or catch me at the booth!
We will present ROVER at #ICLR2026. Our poster session is at Pavilion 4 P4-#3016 on Fri, Apr 24th, 3:15 p.m. - 5:45 p.m. BRT. If you're interested in unified multimodal models, omnimodal generation and cross-modal reasoning, please check it out! I'd be happy to discuss our work remotely🥰
Thrilled to share that I’ve joined @GoogleDeepMind as a Research Scientist.
Excited for what’s ahead and the amazing people I’ll get to work with. 🚀
I am based in Kirkland, happy to meet old and new friends here (Feel free to ping me over WeChat or email)!
Unified multimodal models can generate text and images, but can they truly reason across modalities? 🎨
Introducing ROVER, the first benchmark that evaluates reciprocal cross-modal reasoning in unified models, the next frontier of omnimodal intelligence.
🌐 Project: https://t.co/qA5EPaK5s7
📄 Paper: https://t.co/UjLGs3ZFel
📂 Benchmark: https://t.co/2oyk8SyYOo
🚀 BAGEL — the Unified Multimodal Model with emergent capabilities and production-ready performance — is finally live!
Dive in here:
👉 https://t.co/smEMEK1jMn