Microsoft's MAI-Image-2.5-Pro debuts at #1 on the Artificial Analysis Image Editing Leaderboard, and takes the #7 spot in Text to Image
MAI-Image-2.5-Pro is Microsoft AI's quality-focused image model, launched July 23 in preview on Microsoft Foundry. It joins MAI-Image-2.5 and MAI-Image-2.5-Flash in a family Microsoft positions as covering the quality-speed-cost curve, so builders can pick the point that fits their job. Pro sits at the quality end: Microsoft describes it as its highest-fidelity image model to date, aimed at hero imagery, detailed editing, and precise in-image text rendering.
In the Artificial Analysis Image Arena, MAI-Image-2.5-Pro debuts at #1 on the Image Editing Leaderboard, surpassing Reve 2.1, GPT Image 2, and Microsoft's own MAI-Image-2.5. In Text to Image it lands at #7.
On Microsoft Foundry, MAI-Image-2.5-Pro is priced per token: $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens, which works out to roughly $108.5 per 1k 1024x1024 images. That compares to $48 per 1k images for MAI-Image-2.5 and $20 per 1k for MAI-Image-2.5-Flash.
MAI-Image-2.5-Pro is available in preview on Microsoft Foundry across seven global-standard regions, and can be tried out in the MAI Playground.
Congratulations to @MicrosoftAI on the release!
See below for comparisons between MAI-Image-2.5-Pro and other leading models in the Artificial Analysis Image Arena 🧵
MAI-Image-2.6 is now the #2 text-to-image model in the world - beating out Nano Banana, Meta, and Grok! Fantastic moment for the team who've been hill climbing relentlessly. Try it out now on Arena!
Excited to share that our first MAI image editing model, MAI-Image-2.5, is now ranked #2 on LMArena’s image editing leaderboard 🚀🚀🚀
It has been an amazing journey building this model from scratch alongside old friends and new teammates!
MAI-Image-2.5 has officially released from @MicrosoftAI landing at #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401 and advances the Pareto frontier!
This puts the model +10 pts over Nano Banana 2, Grok Imagine Image Quality and ChatGPT-Image-Latest-High Fidelity.
Congrats to the @MicrosoftAI team on this big accomplishment!
Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier.
First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks.
- It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities.
- It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks.
- And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end.
Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing.
Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI.
- Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost.
All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat.
Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost.
Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare.
Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: https://t.co/v65eop5Ixq
Text is often the hardest part of image generation to get right. MAI-Image-2 improves consistency and legibility for in-image text across infographics, diagrams, and slides — reducing the gap between prompt and output. Try it for yourself. https://t.co/Q8PNLJyZCU
Three models. Three top-tier results. All shipped within just a few months by the @MicrosoftAI team.
- MAI-Transcribe-1 dropped today, the most accurate transcription model in the world across 25 languages according to FLEURS WER benchmark.
- MAI-Voice-1 sets a new standard for natural speech.
- MAI-Image-2 lands as a top 3 model family on @arena.
We've been building with them - now you can too. All 3 available now on Microsoft Foundry.
🚀 Introducing our fresh work at Stanford and Meta MSL:
UniT — Unified Multimodal Chain-of-Thought Test-time Scaling
What if a single model could generate an image, look at it, think about what's wrong, and fix it — all by itself?
That's exactly what UniT does. 🧵👇
Meta just announced MoCha
This AI can create full movie-quality talking & singing characters from just speech & text.
10 wild examples:
1. Talking Characters
Excited to share our latest work, MoCha☕️, led by our incredible intern @CongWei1230! MoCha takes talking-character generation to the next level—unlike previous I2V-based methods, it directly generates movie-grade single- or multi-character videos from raw speech and text input!
🚀Thrilled to introduce ☕️MoCha: Towards Movie-Grade Talking Character Synthesis
Please unmute to hear the demo audio.
✨We defined a novel task: Talking Characters, which aims to generate character animations directly from Natural Language and Speech input.
✨We propose MoCha, the first-of-its-kind DiT model capable of achieving movie-grade talking character generation.
✨MoCha enables, for the first time, Multi-character Conversations with Turn-based Dialogue generation, pushing the boundaries of automated filmmaking.
Paper: https://t.co/1fJfGEFavi
Project website: https://t.co/I1XUY1GLUz
🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date.
Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike.
More details and examples of what Movie Gen can do ➡️ https://t.co/M19x2ndwnr
🛠️ Movie Gen models and capabilities
Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt.
Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment.
Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes.
Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video.
We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.
Meta presents Imagine yourself
Tuning-Free Personalized Image Generation
paper page: https://t.co/WVTywILWfH
Diffusion models have demonstrated remarkable efficacy across various image-to-image tasks. In this research, we introduce Imagine yourself, a state-of-the-art model designed for personalized image generation. Unlike conventional tuning-based personalization techniques, Imagine yourself operates as a tuning-free model, enabling all users to leverage a shared framework without individualized adjustments. Moreover, previous work met challenges balancing identity preservation, following complex prompts and preserving good visual quality, resulting in models having strong copy-paste effect of the reference images. Thus, they can hardly generate images following prompts that require significant changes to the reference image, \eg, changing facial expression, head and body poses, and the diversity of the generated images is low. To address these limitations, our proposed method introduces 1) a new synthetic paired data generation mechanism to encourage image diversity, 2) a fully parallel attention architecture with three text encoders and a fully trainable vision encoder to improve the text faithfulness, and 3) a novel coarse-to-fine multi-stage finetuning methodology that gradually pushes the boundary of visual quality. Our study demonstrates that Imagine yourself surpasses the state-of-the-art personalization model, exhibiting superior capabilities in identity preservation, visual quality, and text alignment. This model establishes a robust foundation for various personalization applications. Human evaluation results validate the model's SOTA superiority across all aspects (identity preservation, text faithfulness, and visual appeal) compared to the previous personalization models.
Excited to share the first project I’ve worked on since I joined GenAI, Meta. I am very proud to work with such an amazing team. Try it in IG, Messenger, and https://t.co/GGvYUqpLLk.
🆕 Research paper from GenAI at Meta: Imagine yourself: Tuning-Free Personalized Image Generation.
Research paper ➡️ https://t.co/8RlWdU5MKu
Want to try it? The feature is available now as a beta in Meta AI for users in the US.
We put a very (very!) fun thing in Meta AI today.
Say "imagine me..." to see yourself anywhere your heart desires. If you're weird like me, you might imagine yourself with a magical emu.
🔜 try it in IG, Messenger, and https://t.co/QYLz1R3Vzh — @AIatMeta
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
paper page: https://t.co/JECZdMqwyc
Recent advances in generative AI have significantly enhanced image and video editing, particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However, the computational demands of diffusion-based methods are substantial, often necessitating large-scale paired datasets for training, and therefore challenging the deployment in practical applications. This study addresses this challenge by breaking down the text-based video editing process into two separate stages. In the first stage, we leverage an existing text-to-image diffusion model to simultaneously edit a few keyframes without additional fine-tuning. In the second stage, we introduce an efficient model called MaskINT, which is built on non-autoregressive masked generative transformers and specializes in frame interpolation between the keyframes, benefiting from structural guidance provided by intermediate frames. Our comprehensive set of experiments illustrates the efficacy and efficiency of MaskINT when compared to other diffusion-based methodologies. This research offers a practical solution for text-based video editing and showcases the potential of non-autoregressive masked generative transformers in this domain.