Excited to share our progress on Movie Gen, a SOTA model for video generation! 🎥✨
I worked on this project as part of a cutting-edge team 🔥, pushing the boundaries of video editing ✂️— all without supervised data.
Can’t wait to show you what’s next! 🚀🎬
🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date.
Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike.
More details and examples of what Movie Gen can do ➡️ https://t.co/M19x2ndwnr
🛠️ Movie Gen models and capabilities
Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt.
Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment.
Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes.
Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video.
We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.
Today we launched Muse Image and are previewing Muse Video! 🚀
https://t.co/sVPBIHGgzv 👈
Muse Image does agentic image generation! It's great at tool use and search. I highly recommend trying it out on https://t.co/JBfhbJUJLB!
Excited to continue! 🧑🍳🔥
Glad to see Muse Image launched and perform so well on @arena ! Multi-image editing was the last thing I worked on at Meta, with the amazing Muse Image team. Looking forward to what folks will create with it! 🥭
Exciting news: Meta’s Muse Image just claimed #2 in the Image Arena!
Muse Image from @AIatMeta now ranks second only to OpenAI's GPT Image 2, outperforming Nano Banana, Grok Imagine, MAI Image, and many other leading image models.
It holds #2 across the board: Text-to-Image, Single-Image Edit, and Multi-Image Edit. Congrats to the Meta team on this incredible milestone!
We are excited to launch muse Image, the first media generation model built by MSL. Muse image can use reasoning, refinement, and tools to improve precision and quality, with very clear test time scaling trends. Also sharing Muse Video preview today.
https://t.co/GOHEd6GWIJ
We’re launching Muse Image today!! 🎉 It’s an agentic image gen model that plans, writes code and uses search tools, and refines its own outputs in chain-of-thought. Image performance improves as we scale test-time compute with higher reasoning. https://t.co/LUuvpixsui
Introducing Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs.
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Muse Spark is available today at https://t.co/wHkMPH82ZH and the Meta AI app. We’re also making it available in private preview via API to select partners, and we hope to open-source future versions of the model.
Learn more: https://t.co/PloE9q5x96
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵
Video models as Physics simulators. 🌍🎥
[1/] In our latest work, WinDiNet, we finetuned a pre-trained video model into a differentiable physics engine. 1000x faster than traditional CFD solvers.
Project page: https://t.co/LAx7t00y3e
Abs: https://t.co/OdcgbKeQEG
**Transition Matching** is a new iterative generative paradigm using Flow Matching or AR models to transition between generation intermediate states, leading to an improved generation quality and speed!
Introducing our first set of Llama 4 models!
We’ve been hard at work doing a complete re-design of the Llama series. I’m so excited to share it with the world today and mark another major milestone for the Llama herd as we release the *first* open source models in the Llama 4 collection 🦙. Here are some highlights:
📌 The Llama series have been re-designed to use state of the art mixture-of-experts (MoE) architecture and natively trained with multimodality. We’re dropping Llama 4 Scout & Llama 4 Maverick, and previewing Llama 4 Behemoth.
📌 Llama 4 Scout is highest performing small model with 17B activated parameters with 16 experts. It’s crazy fast, natively multimodal, and very smart. It achieves an industry leading 10M+ token context window and can also run on a single GPU!
📌 Llama 4 Maverick is the best multimodal model in its class, beating GPT-4o and Gemini 2.0 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeek v3 on reasoning and coding – at less than half the active parameters. It offers a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena. It can also run on a single host!
📌 Previewing Llama 4 Behemoth, our most powerful model yet and among the world’s smartest LLMs. Llama 4 Behemoth outperforms GPT4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.
A big thanks to all of our launch partners (full list in blog) for helping us bring Llama 4 to developers everywhere including @huggingface, @togethercompute, @SnowflakeDB, @ollama, @databricks and many others👏 This is just the start, we have more models coming and the team is really cooking – look out for Llama 4 Reasoning 😉
A few weeks ago, we celebrated Llama being downloaded over 1 billion times. Llama 4 demonstrates our long-term commitment to open source AI, the entire open source AI community, and our unwavering belief that open systems will produce the best small, mid-size and soon frontier models. Llama would be nothing without the global open source AI community & we are so ready to begin this next chapter with you. 🦙
Read more about the release here: https://t.co/7mbK3uggjO, and try it in our products today.
I'm thrilled to announce that Through-The-Mask (TTM) has been accepted to #CVPR2025!
TTM is an I2V generation framework that leverages mask-based motion trajectories to enhance object-specific motion and maintain consistency, especially in multi-object scenarios
More details👇
🚀 Introducing VideoJAM – a framework that instills a strong motion prior into any video model! By denoising an optical flow derivative alongside pixels, VideoJAM teaches models to generate coherent motion and physics with high-quality visuals. 📽️
VideoJAM is our new framework for improved motion generation from @AIatMeta
We show that video generators struggle with motion because the training objective favors appearance over dynamics.
VideoJAM directly adresses this **without any extra data or scaling**
👇🧵
This is extremely cool!
They find diffusion loss is not very sensitive to motion. Thus they fine-tune videogen models with additional explicit motion prediction, making the model generate much more coherent videos.
Also, Hila has been doing consistently good work, follow her!
Meta just dropped VideoJAM
Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
comparison with openai sora and kling
VideoJAM is our new framework for improved motion generation from @AIatMeta
We show that video generators struggle with motion because the training objective favors appearance over dynamics.
VideoJAM directly adresses this **without any extra data or scaling**
👇🧵
Great work on image-to-video generation led by the amazing @guy_yariv during his internship with our team 🖼️➡️👤➡️🎥
Paper: https://t.co/f46Y9mlIor
page: https://t.co/A2WqLCYjeA
[1/8] Recent work has shown impressive Image-to-Video (I2V) generation results. However, accurately articulating multiple interacting objects and complex motions remains challenging. In our new work, we take a step toward addressing this challenge.
[1/8] Recent work has shown impressive Image-to-Video (I2V) generation results. However, accurately articulating multiple interacting objects and complex motions remains challenging. In our new work, we take a step toward addressing this challenge.
VERY excited about the era of generative AR we're bringing to life. Check out this preview!
It's early but so damn promising — this isn't "AI slop"... it's unlocking Creators' imaginations on their own videos. Change your wardrobe, scene, lighting etc. with little expertise.
PS it's been so damn special to navigate this idea maze with some of the best & brightest folks from all across Meta. A highlight of my time here so far.