We feel your pain with video generation glitches: face drift, melting hands, outfit changes, and physics that looks fine until the second watch.
AI video is getting magical — but production doesn’t happen at first glance. It happens through rewatches, edits, reviews, clients, directors, and audiences, where every frame has to hold up.
That’s why we’re opening the interactive demo for 𝐏𝐡𝐲𝐬𝐢𝐨𝐧-𝐀𝐭𝐥𝐚𝐬 1.0: to make video generation failures visible, inspectable, and grounded in real evidence.
You can interact with the demo yourself: inspect videos, reveal hidden glitches, and see how many failures you can successfully spot.
- Interactive demo: https://t.co/LPT6U5MQ6J
- Blog: https://t.co/1Ck3CpuJk6
We show an apples-to-apples view across 𝐒𝐞𝐞𝐝𝐚𝐧𝐜𝐞 2.0, 𝐕𝐞𝐨 3.1, 𝐊𝐥𝐢𝐧𝐠 3.0, 𝐇𝐚𝐩𝐩𝐲𝐡𝐨𝐫𝐬𝐞 1.0, 𝐇𝐚𝐩𝐩𝐲𝐁𝐚𝐧𝐚𝐧𝐚, 𝐏𝐢𝐱𝐯𝐞𝐫𝐬𝐞 𝐕6, and 𝐆𝐫𝐨𝐤 𝐈𝐦𝐚𝐠𝐢𝐧𝐞 — surfacing the subtle, production-critical issues that are easy to miss but hard to ignore.
Come see where generated worlds start to break 👀
Today, we’re introducing 𝐏𝐡𝐲𝐬𝐢𝐨𝐧 𝐋𝐞𝐚𝐝𝐞𝐫𝐛𝐨𝐚𝐫𝐝𝐬: https://t.co/rRJhpqiipb🚀🚀🚀
Over the past year, we’ve spent a lot of time looking closely at how video models fail—not just whether a clip looks impressive, but whether characters remain consistent, objects persist, actions have plausible consequences, and longer videos hold together.
That work has made one thing clear: the field needs evaluation that is more rigorous, more diagnostic, and more honest about what these systems can and cannot do today.
Physion Leaderboards is our commitment to building that together with the community.
The goal is not simply to produce another ranking. We want to show where a model is genuinely strong, where it remains fragile, and what progress is actually happening underneath the demos.
Our evaluations combine careful human judgment, detailed rubrics, and evidence from the videos themselves across physical realism, temporal consistency, identity, causality, narrative, cinematic language, and production quality. Over time, we hope to turn this expertise into reliable automated critic capabilities, as discussed in our Galileo-0 work (https://t.co/fJ0Te3Wbk6). 𝐁𝐮𝐭 𝐰𝐞 𝐚𝐫𝐞 𝐬𝐭𝐢𝐥𝐥 𝐞𝐚𝐫𝐥𝐲. 𝐖𝐞 𝐰𝐚𝐧𝐭 𝐭𝐨 𝐮𝐧𝐝𝐞𝐫𝐬𝐭𝐚𝐧𝐝 𝐰𝐡𝐚𝐭 𝐠𝐨𝐨𝐝 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧 𝐥𝐨𝐨𝐤𝐬 𝐥𝐢𝐤𝐞 𝐟𝐢𝐫𝐬𝐭—𝐚𝐧𝐝 𝐚𝐮𝐭𝐨𝐦𝐚𝐭𝐞 𝐨𝐧𝐥𝐲 𝐰𝐡𝐚𝐭 𝐰𝐞 𝐜𝐚𝐧 𝐯𝐚𝐥𝐢𝐝𝐚𝐭𝐞 𝐫𝐞𝐬𝐩𝐨𝐧𝐬𝐢𝐛𝐥𝐲.
We hope Physion Leaderboards can become a useful shared reference for the community, and over time help contribute to an evaluation standard that is rigorous, adaptable, and independent enough to reflect the frontier as faithfully as possible ✨
Explore the first two leaderboards:
- Physion-Atlas 1.0: https://t.co/8KzVp4oDKU
- Physion-Arc 1.0: https://t.co/EPuzlidlD8
🎬 Video agents can now generate minute-long videos. But can they actually direct?
Today, we’re launching 𝐏𝐡𝐲𝐬𝐢𝐨𝐧-𝐀𝐫𝐜 1.0, a new benchmark evaluating complete, multi-scene videos across narrative coherence, cinematic language, and production quality.
We tested @runwayml, @LumaLabsAI, @MiniMax_AI, @Kling_ai , @UtopaiStudios and @TapNow_AI on 100 screenplays and 600 generated videos.
🏆 𝐑𝐮𝐧𝐰𝐚𝐲 𝐀𝐠����𝐧𝐭 2.0 𝐫𝐚𝐧𝐤𝐞𝐝 𝐍𝐨. 1 𝐨𝐯𝐞𝐫𝐚𝐥𝐥 𝐚𝐧𝐝 𝐥𝐞𝐝 𝐞𝐯𝐞𝐫𝐲 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧 𝐝𝐢𝐦𝐞𝐧𝐬𝐢𝐨𝐧. Its advantage was especially clear across subjective metrics, where cinematic taste matters most. Runway ranked first on all eight.
🔗 Read the full benchmark: https://t.co/PuBpmwUQGv
We've spoken with hundreds of ad creatives, marketing designers, filmmakers, and animation teams — and heard the same thing: the outputs look great… until they don't 😅. When they fail, it's incredibly hard to tell why. Is it the prompt, the model, or the world itself quietly breaking? That ambiguity is the real bottleneck.
Physion-Atlas 1.0 introduces a more objective, diagnostic way to evaluate video world models — moving beyond high-level comparisons to surface what actually matters. It disentangles prompt misalignment from physical and visual inconsistencies, grounding every judgment in explicit spatiotemporal evidence. Not just which output is better, but what breaks, when, where, and why. From abstract comparisons → diagnosable reality 🔍
📄 Blog: https://t.co/9aKSKfR7yS
📝 Evaluate your model: https://t.co/Jm1fshiF4a
We’ve been quietly blown away by the response to Galileo-0: https://t.co/fJ0Te3Wbk6
Over the past few days, we’ve seen waiting list signups from a wide range of teams — from video generation startups and frontier model developers to video platforms and studios.
We’re grateful for the curiosity and openness to engage. It’s still early, and we have a lot to learn.
Over the coming week, we’ll start reaching out and connecting with teams on the waiting list to better understand your workflows, challenges, and where a world critic like Galileo-0 can be most useful.
If you’ve signed up — thank you. Looking forward to the conversations ahead.
#Galileo0 #PhysionLabs #WorldCritic
🚀🚀🚀We're excited to introduce Galileo 0 (https://t.co/rWeqZzMTx1) — our first research preview of a world critic for AI-generated video, which already outperforms Qwen 3.5-Plus, Gemini 3.1 Pro, Pegasus 1.2, and GPT 5.4 on physical consistency reasoning 🚀🚀🚀
Galileo doesn't just score outputs. It diagnoses them — identifying what failed, when it failed, where it happened, and why it broke the rules of the world.
This is a step toward a new paradigm: generate → critique → refine → repeat — where models don't just produce worlds, but learn to keep them consistent over time.
𝐖𝐡𝐚𝐭 𝐦𝐚𝐤𝐞𝐬 𝐭𝐡𝐢𝐬 𝐦𝐢𝐥𝐞𝐬𝐭𝐨𝐧𝐞 𝐞𝐯𝐞𝐧 𝐦𝐨𝐫𝐞 𝐦𝐞𝐚𝐧𝐢𝐧𝐠𝐟𝐮𝐥:
We built Galileo 0 — along with our datasets (including our public Physion-Eval benchmark), evaluation pipeline, and early pilots — with less than $200K total spend in 3 months.
No massive training clusters. No billion-dollar budgets. Just a small, relentless team, strong conviction — and yes, at one point, a five-day stretch of not showering to get this model out.
Because we believe reliability will become core infrastructure for world models.
In a world where billions are being poured into generation, the missing piece isn't more pixels — it's better critics 😊
#PhysionLabs #Galileo0 #WorldModels
🚨 “𝐖𝐡��𝐜𝐡 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐞𝐝 𝐯𝐢𝐝𝐞𝐨 𝐥𝐨𝐨𝐤𝐬 𝐛𝐞𝐭𝐭𝐞𝐫?” is the wrong question. And yet, that’s exactly what arena-style evaluation asks — and much of multimodal AI is still judged this way. The problem is that it captures visual preference, but fails to measure whether the scene is actually coherent — whether objects behave consistently, interactions make sense, or events follow causal structure.
The real challenge isn’t visual quality. It’s whether a model can produce outputs that are correctly grounded across space, time, objects, and interactions — in other words, 𝐝𝐞𝐭𝐚𝐢𝐥𝐞𝐝 𝐦𝐮𝐥𝐭𝐢𝐦𝐨𝐝𝐚𝐥 𝐠𝐫𝐨𝐮𝐧𝐝𝐢𝐧𝐠 𝐚𝐧𝐝 𝐫𝐞𝐚𝐬𝐨𝐧𝐢𝐧𝐠.
At 🚀 𝐏𝐡𝐲𝐬𝐢𝐨𝐧 𝐋𝐚𝐛𝐬 🚀, in collaboration with researchers from 𝐒𝐭𝐚𝐧𝐟𝐨𝐫𝐝, 𝐌𝐈𝐓, and 𝐇𝐚𝐫𝐯𝐚𝐫𝐝 -- including Peiyu Jing, Hong-Xing "Koven" Yu, Fangqiang Ding, Fan Nie, Weimin Wang, Yilun Du, James Zou, Jiajun Wu, and Bing Shuai -- we analyzed state-of-the-art video generation models. What we found is hard to ignore: 𝐚𝐜𝐫𝐨𝐬𝐬 𝐥𝐞𝐚𝐝𝐢𝐧𝐠 𝐯𝐢𝐝𝐞𝐨 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬, 𝟴𝟯.𝟯% 𝐨𝐟 𝐞𝐱𝐨𝐜𝐞𝐧𝐭𝐫𝐢𝐜 𝐯𝐢𝐝𝐞𝐨𝐬 𝐚𝐧𝐝 𝟵𝟯.𝟱% 𝐨𝐟 𝐞𝐠𝐨𝐜𝐞𝐧𝐭𝐫𝐢𝐜 𝐯𝐢𝐝𝐞𝐨𝐬 𝐜𝐨𝐧𝐭𝐚𝐢𝐧 𝐩𝐡𝐲𝐬𝐢𝐜𝐚𝐥 𝐢𝐧𝐜𝐨𝐧𝐬𝐢𝐬𝐭𝐞𝐧𝐜𝐢𝐞𝐬. These are not just visual artifacts, but failures in object interactions, temporal continuity, and causal structure. Many are subtle, but fundamentally wrong.
This reveals a critical gap. We’ve made massive progress in making videos look better, but far less progress in making them actually grounded and consistent. The uncomfortable truth is that “looks right” does not mean “is right,” and preference does not imply understanding.
We’re releasing 🎬 𝐏𝐇𝐘��𝐈𝐎𝐍-𝐄𝐕𝐀𝐋, the first human-centered benchmark for physical realism in AI-generated video. It includes over 10,000 expert reasoning traces, spans 22 fine-grained physical phenomena, provides temporally grounded annotations, and enables direct comparison between human and model reasoning.
📄 Paper: https://t.co/uIwSYBHva1
��� Dataset: https://t.co/QDeWxi26gU
🖼️ Preview: https://t.co/fUUrWj5XZD
𝐈𝐟 𝐰𝐞 𝐤𝐞𝐞𝐩 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐢𝐧𝐠 𝐟𝐨𝐫 𝐚𝐩𝐩𝐞𝐚𝐫𝐚𝐧𝐜𝐞, 𝐰𝐞’𝐥𝐥 𝐠𝐞𝐭 𝐦𝐨𝐫𝐞 𝐜𝐨𝐧𝐯𝐢𝐧𝐜𝐢𝐧𝐠 𝐢𝐥𝐥𝐮𝐬𝐢𝐨𝐧𝐬 — 𝐧𝐨𝐭 𝐦𝐨𝐫𝐞 𝐫𝐞𝐥𝐢𝐚𝐛𝐥𝐞 𝐬𝐲𝐬𝐭𝐞𝐦𝐬. And for world models, robotics, and real-world deployment, that’s a fundamental failure.
We’re open-sourcing the dataset and releasing the paper today. This is a step toward a new standard: not just generating what looks good, but generating what is actually 𝐠𝐫𝐨𝐮𝐧𝐝𝐞𝐝, 𝐜𝐨𝐧𝐬𝐢𝐬𝐭𝐞𝐧𝐭, 𝐚𝐧𝐝 𝐜𝐨𝐫𝐫𝐞𝐜𝐭.
#AI #VideoGeneration #MultimodalAI #AIEvaluation #AIBenchmark #WorldModels #DeepLearning 🚀🐶