๐จ โ๐๐ก๐ข๐๐ก ๐ ๐๐ง๐๐ซ๐๐ญ๐๐ ๐ฏ๐ข๐๐๐จ ๐ฅ๐จ๐จ๐ค๐ฌ ๐๐๐ญ๐ญ๐๐ซ?โ is the wrong question. And yet, thatโs exactly what arena-style evaluation asks โ and much of multimodal AI is still judged this way. The problem is that it captures visual preference, but fails to measure whether the scene is actually coherent โ whether objects behave consistently, interactions make sense, or events follow causal structure.
The real challenge isnโt visual quality. Itโs whether a model can produce outputs that are correctly grounded across space, time, objects, and interactions โ in other words, ๐๐๐ญ๐๐ข๐ฅ๐๐ ๐ฆ๐ฎ๐ฅ๐ญ๐ข๐ฆ๐จ๐๐๐ฅ ๐ ๐ซ๐จ๐ฎ๐ง๐๐ข๐ง๐ ๐๐ง๐ ๐ซ๐๐๐ฌ๐จ๐ง๐ข๐ง๐ .
At ๐ ๐๐ก๐ฒ๐ฌ๐ข๐จ๐ง ๐๐๐๐ฌ ๐, in collaboration with researchers from ๐๐ญ๐๐ง๐๐จ๐ซ๐, ๐๐๐, and ๐๐๐ซ๐ฏ๐๐ซ๐ -- including Peiyu Jing, Hong-Xing "Koven" Yu, Fangqiang Ding, Fan Nie, Weimin Wang, Yilun Du, James Zou, Jiajun Wu, and Bing Shuai -- we analyzed state-of-the-art video generation models. What we found is hard to ignore: ๐๐๐ซ๐จ๐ฌ๐ฌ ๐ฅ๐๐๐๐ข๐ง๐ ๐ฏ๐ข๐๐๐จ ๐ ๐๐ง๐๐ซ๐๐ญ๐ข๐จ๐ง ๐ฆ๐จ๐๐๐ฅ๐ฌ, ๐ด๐ฏ.๐ฏ% ๐จ๐ ๐๐ฑ๐จ๐๐๐ง๐ญ๐ซ๐ข๐ ๐ฏ๐ข๐๐๐จ๐ฌ ๐๐ง๐ ๐ต๐ฏ.๐ฑ% ๐จ๐ ๐๐ ๐จ๐๐๐ง๐ญ๐ซ๐ข๐ ๐ฏ๐ข๐๐๐จ๐ฌ ๐๐จ๐ง๐ญ๐๐ข๐ง ๐ฉ๐ก๐ฒ๐ฌ๐ข๐๐๐ฅ ๐ข๐ง๐๐จ๐ง๐ฌ๐ข๐ฌ๐ญ๐๐ง๐๐ข๐๐ฌ. These are not just visual artifacts, but failures in object interactions, temporal continuity, and causal structure. Many are subtle, but fundamentally wrong.
This reveals a critical gap. Weโve made massive progress in making videos look better, but far less progress in making them actually grounded and consistent. The uncomfortable truth is that โlooks rightโ does not mean โis right,โ and preference does not imply understanding.
Weโre releasing ๐ฌ ๐๐๐๐๐๐๐-๐๐๐๐, the first human-centered benchmark for physical realism in AI-generated video. It includes over 10,000 expert reasoning traces, spans 22 fine-grained physical phenomena, provides temporally grounded annotations, and enables direct comparison between human and model reasoning.
๐ Paper: https://t.co/uIwSYBHva1
๐ค Dataset: https://t.co/QDeWxi26gU
๐ผ๏ธ Preview: https://t.co/fUUrWj5XZD
๐๐ ๐ฐ๐ ๐ค๐๐๐ฉ ๐จ๐ฉ๐ญ๐ข๐ฆ๐ข๐ณ๐ข๐ง๐ ๐๐จ๐ซ ๐๐ฉ๐ฉ๐๐๐ซ๐๐ง๐๐, ๐ฐ๐โ๐ฅ๐ฅ ๐ ๐๐ญ ๐ฆ๐จ๐ซ๐ ๐๐จ๐ง๐ฏ๐ข๐ง๐๐ข๐ง๐ ๐ข๐ฅ๐ฅ๐ฎ๐ฌ๐ข๐จ๐ง๐ฌ โ ๐ง๐จ๐ญ ๐ฆ๐จ๐ซ๐ ๐ซ๐๐ฅ๐ข๐๐๐ฅ๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ๐ฌ. And for world models, robotics, and real-world deployment, thatโs a fundamental failure.
Weโre open-sourcing the dataset and releasing the paper today. This is a step toward a new standard: not just generating what looks good, but generating what is actually ๐ ๐ซ๐จ๐ฎ๐ง๐๐๐, ๐๐จ๐ง๐ฌ๐ข๐ฌ๐ญ๐๐ง๐ญ, ๐๐ง๐ ๐๐จ๐ซ๐ซ๐๐๐ญ.
#AI #VideoGeneration #MultimodalAI #AIEvaluation #AIBenchmark #WorldModels #DeepLearning ๐๐ถ
๐ข RadarGen: Automotive Radar Point Cloud Generation from Cameras
Can we generate realistic radar point clouds solely from camera images? ๐๐ก
We introduce RadarGen, a diffusion-based framework that synthesizes radar returns aligned with visual scenes.
https://t.co/IYCQ8mVcmC
If you are attending #CVPR2023 physically or virtually, please check out ๐our Highlight ๐ก paper "Hidden Gems๐: 4D Radar Scene Flow Learning Using Cross-Modal Supervision" (WED-AM-106) (co-authored with @chris_x_lu@andraspalffy and Dariu Gavrila). Page: https://t.co/0hkzL1eLBf
I will be at @ieee_ras_icra ICRA2023 next week, presenting my poster on Tue 30 May 15:00-16:40, PODS 1-15.
Look forward to meet many people who are also enthusiastic about robotics. Do take time to drop by and have a chat with me๐.
Exciting news! Our #CVPR2023 paper "Hidden Gems: 4D Radar Scene Flow Learning Using Cross-Modal Supervision" was selected as a highlight ๐ก (10% of the accepted papers, 2.5% of submissions). Can't wait ๐ to share our ideas with every attendee.
๐Thrilled to announce our #CVPR2023 work "๐๐ข๐๐๐๐ง ๐๐๐ฆ๐ฌ: 4๐ ๐๐๐๐๐ซ ๐๐๐๐ง๐ ๐ ๐ฅ๐จ๐ฐ ๐๐๐๐ซ๐ง๐ข๐ง๐ ๐๐ฌ๐ข๐ง๐ ๐๐ซ๐จ๐ฌ๐ฌ-๐๐จ๐๐๐ฅ ๐๐ฎ๐ฉ๐๐ซ๐ฏ๐ข๐ฌ๐ข๐จ๐ง" got accepted. Am approach to 4D radar-based scene flow estimation via cross-modal learning was proposed.
๐Thrilled to announce our #CVPR2023 work "๐๐ข๐๐๐๐ง ๐๐๐ฆ๐ฌ: 4๐ ๐๐๐๐๐ซ ๐๐๐๐ง๐ ๐ ๐ฅ๐จ๐ฐ ๐๐๐๐ซ๐ง๐ข๐ง๐ ๐๐ฌ๐ข๐ง๐ ๐๐ซ๐จ๐ฌ๐ฌ-๐๐จ๐๐๐ฅ ๐๐ฎ๐ฉ๐๐ซ๐ฏ๐ข๐ฌ๐ข๐จ๐ง" got accepted. Am approach to 4D radar-based scene flow estimation via cross-modal learning was proposed.
Glad to see Fangqiang Ding presenting our recent work on self-supervised scene flow estimation with automotive radars. This work was also recently accepted by RA-L/IROS. Well done @Toytiny3!
We have 14 funded places available on our PhD in Robotics and Autonomous Systems programme starting September 2022. More details and how to apply https://t.co/Ydl45RrmJ4