Thrilled to share results from the Movie Gen models we've been working on these past few months, and particularly the Movie Gen Edit model for precise editing! 🚀🚀
🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date.
Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike.
More details and examples of what Movie Gen can do ➡️ https://t.co/M19x2ndwnr
🛠️ Movie Gen models and capabilities
Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt.
Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment.
Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes.
Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video.
We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.
DLSS 5 is all over the timeline, and for good reason. In my internship at @AIatMeta we had the same idea: use a video model as a learned second-stage renderer on top of game engines. In our paper RealMaster, we make synthetic video look real while preserving scene fidelity 👇
New research from @bfl_ml 🥳
Meet Self-Flow: our self-supervised framework for image, audio, video & world models 🤖
https://t.co/AshY8IkSEe
Do generative models really need DINO to learn strong representations? We propose teaching them directly via a joint framework instead 🧵
1/9
Excited to share EditP23! 🎨
Finally, a single tool for ALL your 3D editing needs:
✅ Pose & Geometry Changes
✅ Object Additions
✅ Global Style Transformations
✅ Local Modifications
All driven by one simple 2D image edit. It's mask-free ✨ and works in seconds ⚡️.
🧵
[1/n]
New paper alert! 🚀
Excited to introduce 𝐓𝐫𝐚𝐧𝐬𝐢𝐭𝐢𝐨𝐧 𝐌𝐚𝐭𝐜𝐡𝐢𝐧𝐠 (𝐓𝐌)! We're replacing short-timestep kernels from Flow Matching/Diffusion with... a generative model🤯, achieving SOTA text-2-image generation!
@urielsinger@itai_gat@lipmanya
Exciting news from #ICML2025 & #ICCV2025 🥳
- 🥇 VideoJAM accepted as *oral* at #ICML2025 (top 1%)
- Two talks at #ICCV2025
☝️interpretability in the generative era
✌️video customization
- Organizing two #ICCV2025 workshops
☝️structural priors for vision
✌️long video gen
🧵👇
The longer reasoning LLM thinks - the more likely to be correct, right?
Apparently not.
Presenting our paper: “Don’t Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning”.
Link: https://t.co/Zsp3BD0TU5
1/n
I'm thrilled to announce that Through-The-Mask (TTM) has been accepted to #CVPR2025!
TTM is an I2V generation framework that leverages mask-based motion trajectories to enhance object-specific motion and maintain consistency, especially in multi-object scenarios
More details👇
Super excited to share 🧠MLGym 🦾 – the first Gym environment for AI Research Agents 🤖🔬
We introduce MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks.
The key contributions of our work are:
🕹️ Enables the exploration of different training algorithms for AI Research Agents such as RL
🛠️ Provides a flexible evaluation framework that can accommodate different artifacts such as models, algorithms, or predictions
🤖 Allows researchers to evaluate any model without the need to develop a custom agentic harness
🎯 Introduces 13 diverse open-ended AI Research tasks for evaluating AI Research Agents on a wide range of domains such as computer vision, natural language processing, reinforcement learning, game theory, and logical reasoning.
📈 Proposes a new evaluation metric for AI Research Agents
MLGym makes it easy to:
1) Add new tasks
2) Evaluate new models
3) Integrate new agents
Check out a video of the MLGym Agent to see how it performs the full pipeline of idea generation💡, implementation 👩💻, experimentation 👩🔬, and iteration 🔄 to improve on ML tasks.
Huge thanks to the exceptionally talented @deepaknathani11 who led this work and to all the other amazing collaborators who made this possible 🙏🫶🚀
This is extremely cool!
They find diffusion loss is not very sensitive to motion. Thus they fine-tune videogen models with additional explicit motion prediction, making the model generate much more coherent videos.
Also, Hila has been doing consistently good work, follow her!
Meta just dropped VideoJAM
Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
comparison with openai sora and kling
🚀 Our latest work, VideoJAM, introduces a new method to enhance motion in any T2V model, significantly improving its motion and physics.
We also train a DiT model that, combined with VideoJAM, achieves a new SOTA in motion generation! 🔥
https://t.co/WZm5RLfgyX
[1/8] Recent work has shown impressive Image-to-Video (I2V) generation results. However, accurately articulating multiple interacting objects and complex motions remains challenging. In our new work, we take a step toward addressing this challenge.
VERY excited about the era of generative AR we're bringing to life. Check out this preview!
It's early but so damn promising — this isn't "AI slop"... it's unlocking Creators' imaginations on their own videos. Change your wardrobe, scene, lighting etc. with little expertise.
PS it's been so damn special to navigate this idea maze with some of the best & brightest folks from all across Meta. A highlight of my time here so far.
Movie Gen claims to be the state-of-the-art in text-to-video generation, outperforming Sora, Kling, Gen3, and more. But how can you trust the results?
Today, we're releasing 1003 videos and their prompts - no cherry-picking allowed. Our goal? To set a new standard for evaluating these models and bring transparency to the field.
Download our videos and prompts in GitHub:
https://t.co/2jw6GjUtg4