the quality vs cost difference here is pretty interesting, especially with @runwayml and @openart_ai ending up so close on overall quality
really excited to see Physion-Arc Ads & Marketing out in the wild ๐
๐ ๐๐ก๐ฒ๐ฌ๐ข๐จ๐ง-๐๐ซ๐ 1.0 ๐๐๐ฌ & ๐๐๐ซ๐ค๐๐ญ๐ข๐ง๐ ๐๐ซ๐๐๐ค ๐ข๐ฌ ๐ฅ๐ข๐ฏ๐.
For this track, we wanted to test something very practical: can todayโs video agents actually take a real advertising brief and turn it into a strong finished ad?
We worked with advertising and creative professionals to build 100 real-world briefs across text-only, single-reference, and multi-reference tasks, then evaluated each agent across 24 metrics covering brief fidelity, production quality, and ad effectiveness.
The first leaderboard includes @runwayml, @openart_ai, @Creatify_AI, @LumaLabsAI, and @HeyGen.
One result we found especially interesting: Runway and OpenArt AI are very close on overall quality โ 71.1 vs. 69.5, with no statistically significant difference in this round.
๐ฐ At the same time, the agents show quite different cost profiles. OpenArt AI is about $6/video, compared with $37/video for Runway, which makes the qualityโcost tradeoff another interesting dimension to look at alongside overall performance.
Full leaderboard here: https://t.co/S13zr74evo
๐ฌ Video agents can now generate minute-long videos. But can they actually direct?
Today, weโre launching ๐๐ก๐ฒ๐ฌ๐ข๐จ๐ง-๐๐ซ๐ 1.0, a new benchmark evaluating complete, multi-scene videos across narrative coherence, cinematic language, and production quality.
We tested @runwayml, @LumaLabsAI, @MiniMax_AI, @Kling_ai , @UtopaiStudios and @TapNow_AI on 100 screenplays and 600 generated videos.
๐ ๐๐ฎ๐ง๐ฐ๐๐ฒ ๐๐ ๐๐ง๐ญ 2.0 ๐ซ๐๐ง๐ค๐๐ ๐๐จ. 1 ๐จ๐ฏ๐๐ซ๐๐ฅ๐ฅ ๐๐ง๐ ๐ฅ๐๐ ๐๐ฏ๐๐ซ๐ฒ ๐๐ฏ๐๐ฅ๐ฎ๐๐ญ๐ข๐จ๐ง ๐๐ข๐ฆ๐๐ง๐ฌ๐ข๐จ๐ง.ย Its advantage was especially clear across subjective metrics, where cinematic taste matters most. Runway ranked first on all eight.
๐ Read the full benchmark: https://t.co/PuBpmwUQGv
We feel your pain with video generation glitches: face drift, melting hands, outfit changes, and physics that looks fine until the second watch.
AI video is getting magical โ but production doesnโt happen at first glance. It happens through rewatches, edits, reviews, clients, directors, and audiences, where every frame has to hold up.
Thatโs why weโre opening the interactive demo for ๐๐ก๐ฒ๐ฌ๐ข๐จ๐ง-๐๐ญ๐ฅ๐๐ฌ 1.0: to make video generation failures visible, inspectable, and grounded in real evidence.
You can interact with the demo yourself: inspect videos, reveal hidden glitches, and see how many failures you can successfully spot.
- Interactive demo: https://t.co/LPT6U5MQ6J
- Blog: https://t.co/1Ck3CpuJk6
We show an apples-to-apples view across ๐๐๐๐๐๐ง๐๐ 2.0, ๐๐๐จ 3.1, ๐๐ฅ๐ข๐ง๐ 3.0, ๐๐๐ฉ๐ฉ๐ฒ๐ก๐จ๐ซ๐ฌ๐ 1.0, ๐๐๐ฉ๐ฉ๐ฒ๐๐๐ง๐๐ง๐, ๐๐ข๐ฑ๐ฏ๐๐ซ๐ฌ๐ ๐6, and ๐๐ซ๐จ๐ค ๐๐ฆ๐๐ ๐ข๐ง๐ โ surfacing the subtle, production-critical issues that are easy to miss but hard to ignore.
Come see where generated worlds start to break ๐
๐New Paper Alert!๐
Introducing POISE for human silhouette extraction! ๐คโจ Our latest research tackles the tricky problem of occlusions with a novel self-supervised approach. (1/5)
Come check out our work on Controllable Dynamic Multi-task Architectures at #CVPR2022 oral session 3.1.2 (8.30-10.30am) and poster session 3.1 (10.00-12.30pm) tomorrow June 23!
Project page: https://t.co/qkDp5nAaRn
I get a lot of reviews that say my work is not novel and I bet I'm not alone. It's always frustrating because I see novelty where the reviewer doesn't. Rather than rebut every critique, I've written a blog post to help reviewers think about novelty. https://t.co/UXLabOkYcn
Errr ok wow, I am shook by the new ConvMixer architecture
https://t.co/crUMktQ0ig "the first model that achieves the elusive dual goals of 80%+ ImageNet top-1 accuracy while also fitting into a tweet" ๐