QuantiPhy challenges VLMs to reason about the physical world with numerical accuracy, estimating size, velocity, and acceleration directly from video.
👉 https://t.co/PhwqXEu4eF
#NeurIPS2026#VLM#MultimodalAI#PhysicalAI#WorldModels
Whether you're working on VLMs, multimodal reasoning, or world models, we'd love to see what your models can do.
Looking forward to seeing everyone on the leaderboard! 🚀🔥
🚀 The QuantiPhy Challenge @ NeurIPS 2026 is officially LIVE!
Can AI truly reason about physics quantitatively?
Team registration and submissions are now open! 🏆
👉 https://t.co/3fnHC3RuFQ
We're excited to invite the community to tackle one of today's most challenging problems in multimodal AI.
✅ Official NeurIPS 2026 Competition Track ✅ Team registration open ✅ Submission portal live ✅ Official leaderboard
Can you identify AI-generated videos by watching how human figures move in the video?
We introduce 🏃♂️HumanScore💃 — a benchmark for evaluating human motions in generated videos through biomechanics, not just visual realism.
Check out our project page: https://t.co/D9LCSgoAq5
#AI #ComputerVision #VideoGeneration #StanfordAI
Current models struggle with basic physics estimates, limiting their use in robotics and autonomous systems. Stanford scholars developed QuantiPhy, a benchmark to evaluate and improve AI's ability to reason about physical properties: https://t.co/4hPbCbD1yL
‼️VLMs/MLLMs do NOT yet understand the physical world from videos‼️
In our recent work, we found that even the most advanced AI models still lag behind humans in one key aspect: reasoning about the kinematic properties of objects from videos.
Takeaways:
1. ChatGPT 5.1 leads overall among 21 advanced VLMs, followed by Gemini 2.5 Pro/Flash.
2. Grok 4.1 delivers impressive performance at the lowest API cost.
3. Qwen3-VL is the top-performing open-source model.
Read here: https://t.co/5lagvLNE37
🧵1/N