Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon.
We’re excited to introduce Muse Spark 1.1, a significant upgrade from the first Muse Spark model we released earlier this year.
Along with this release, we are launching a public preview of the new Meta Model API where developers can access Muse Spark 1.1.
The model is also available now in "Thinking" mode in the Meta AI app and on https://t.co/wHkMPH82ZH.
Learn more: https://t.co/zGcA3XaWpN
Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights
Muse Spark is a new model from @Meta evaluated on Artificial Analysis. We were given early access by Meta to independently benchmark the model. It is the first frontier-class model from Meta since Llama 4 Maverick was released in April 2025, and notably the first @AIatMeta model that is not being released as open weights. The release follows Meta's reorganization of its AI efforts under Meta Superintelligence Labs, and signals that Meta is re-entering the frontier race after roughly a year of relative quiet.
For context, Llama 4 Maverick and Scout scored 18 and 13 respectively on the Artificial Analysis Intelligence Index as non-reasoning models at the time of their release, while Muse Spark scores 52. Muse Spark essentially closes the gap between to the frontier in a single release.
The model is not open source and is not yet accessible via an API but Meta has shared they expect this to come soon. Meta is also integrating Muse Spark into their first party products including their Meta AI chat product, Facebook, Instagram and Threads.
Key takeaways from our benchmarks:
➤ Muse Spark scores 52 on the Artificial Analysis Intelligence Index, placing it within the top 5 models we have benchmarked. It sits ahead of Claude Sonnet 4.6, GLM-5.1, MiniMax-M2.7, Grok 4.20 and behind Gemini 3.1 Pro Preview, GPT-5.4 and Claude Opus 4.6
➤ Muse Spark is notably token efficient for its intelligence level. It used 58M output tokens to run the Intelligence Index, comparable to Gemini 3.1 Pro Preview (57M) and notably lower than Claude Opus 4.6 (Adaptive Reasoning, max effort, 157M), GPT-5.4 (xhigh, 120M) and GLM-5 (110M)
➤ Muse Spark is the second-most capable vision model we have benchmarked. It scores 80.5% on MMMU-Pro, behind only Gemini 3.1 Pro Preview (82.4%)
➤ Muse Spark performs strongly on reasoning and instruction-following evaluations. It scores 39.9% on HLE, trailing only Gemini 3.1 Pro Preview (44.7%) and GPT-5.4 (xhigh, 41.6%). The model also achieved 5th highest in CritPT with a score of 11%, an eval that is focused on difficult physics research questions. This is substantially above above Gemini 3 Flash (9%) and Claude 4.6 Sonnet (3%)
➤ Agentic performance does not stand out. On GDPval-AA, our evalaution focused on real world work tasks, Muse Spark scores 1427, behind both Claude Sonnet 4.6 at 1648 and GPT-5.4 at 1676, but ahead of Gemini 3.1 Pro Preview at 1320. On On TerminalBench Hard, Muse Spark trails Claude Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro. Muse Spark joins others in achieving a high τ²-Bench Telecom score of 92%
Key model details:
➤ Modalities: Multimodal including text and vision input, text output
➤ License: Proprietary, Meta's first frontier model not released as open weights
➤ Availability: No public API at the time of publishing. Meta expects to provide API access soon. Meta has started integration into their first party AI offering Meta AI and inside Facebook, Instagram, and Threads
Introducing Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs.
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Muse Spark is available today at https://t.co/wHkMPH82ZH and the Meta AI app. We’re also making it available in private preview via API to select partners, and we hope to open-source future versions of the model.
Learn more: https://t.co/PloE9q5x96
Excited to launch Viktor
AI COWORKER THAT LIVES IN SLACK
One teammate that handles marketing audits, ad management, lead research, daily reports, and deployed apps. Across every channel. At once.
→ 3,000+ tool integrations. If one's missing, it builds its own
→ Persistent memory. Learns your company, notices patterns, follows up on its own
Today it's yours.
P.S. Early access for now. Slack app review moves at Salesforce speed. Viktor doesn't.
🚨 New Paper: The Art of Scaling Reinforcement Learning Compute for LLMs 🚨
We burnt a lot of GPU-hours to provide the community with the first open, large-scale systematic study on RL scaling for LLMs.
https://t.co/49REQZ4R6G
Today we’re releasing our first public preview of ARC-AGI-3: the first three games.
Version 3 is a big upgrade over v1 and v2 which are designed to challenge pure deep learning and static reasoning. In contrast, v3 challenges interactive reasoning (eg. agents). The full version of v3 will ship early 2026.
ARC v3 games are a bit different than games you’ve played before. Every game is unique and there are no instructions, intentionally. Both humans and AI must play to discover the controls, rules, and goals. Like v1 and v2, only core knowledge is needed to win the games (no language, trivia, etc.)
The core design philosophy of ARC is that it should be relatively easy for humans and hard for AI. V3 takes this even further — it’s the widest gap between humans and AI we’ve shipped thus far.
V3 is our first attempt at creating an interactive benchmark and we’re choosing a different release strategy (today’s preview launch) to get early game design feedback.
In additional to first games, we’re releasing an API that AI researchers can use to test their agents. This is the first piece of real infrastructure that ARC has shipped.
As I’ve previewed the preview with AI researchers over the past few weeks, I’ve felt a higher degree of “pull” for ARC-AGI-3 than I ever felt for v1 and v2.
We hope to steward this attention to continue building useful and interesting benchmarks that guide humanity towards AGI.
Go play the games! And if you’re an AI researcher, check out our API. We’re excited to hear your feedback.
https://t.co/VozET0xRvT hosts live-written fiction—thousands of stories, some millions of words long.
We’re building AI tools to help authors plan, track, and write. But can AI really understand stories that long?
Update: new Grok 3 is solid, LLaMA 4 improves with vLLM fixes 👇
Llama 4 Intelligence Index Update: We have now replicated Meta’s claimed values for MMLU Pro and GPQA Diamond, pushing our Intelligence Index scores for both Scout and Maverick higher
Key update details:
➤ We noted in our first post 48 hours ago that we noticed discrepancies between our measured results and Meta’s claimed scores for our multi-choice eval datasets (MMLU Pro and GPQA Diamond)
➤ After further experiments and and close review, we have decided that in accordance with our published principle against unfairly penalizing models where they get the content of questions correct but format answers differently, we will allow Llama 4’s answer style of ‘The best answer is A’ as legitimate answer for our multi-choice evals
➤ This leads to a jump in score for both Scout and Maverick (largest for Scout) in 2/7 of the evals that make up Artificial Analysis Intelligence Index, and therefore a jump in their Intelligence Index scores
➤ Scout’s Intelligence Index has moved from 36 to 43, and Maverick’s Intelligence Index has moved from 49 to 50.
Overall, we continue to conclude that both Scout and Maverick are very impressive models and a significant contribution to the open weights AI ecosystem.
While DeepSeek V3 0324 maintains a small lead over Maverick, we continue to note that Maverick has ~half the active parameters (17B vs 37B), and ~60% of the total parameters (402B vs 671B), while also supporting image inputs.
All our tests have been performed on the Hugging Face release version of the Llama 4 weights for both Scout and Maverick, including testing via a range of third party cloud providers. None of our eval results are based on the experimental chat-tuned model provided to LMArena (Llama-4-Maverick-03-26-Experimental).
We can also share that we have observed third party cloud APIs generally stabilizing over the last 48 hours. We will soon release endpoint-level comparison data to allow developers to understand whether any cloud providers are still serving versions of Llama 4 with accuracy issues.
🚀 Llama 4 Scout
17B active params, 16 experts, 109B total params
Runs on a single H100 GPU with Int4
10M+ multimodal context for codebases, personalization & video
🚀 Llama 4 Maverick
17B active params, 128 experts, 400B total params
1M+ context
Ranks #2 on LMArena, ELO 1417
Excited to share Llama 4 family of models! 🚀
Our journey to 10M+ multimodal context length has been incredible! The new iRoPE architecture and inference time optimization is a major step toward our long-term goal of infinite context length capabilities🧵
📊 Model Architecture
New iRoPE architecture:
- local 8k attention with RoPE
- interleaved with global NoPE attention layers allows long context generalisation beyond training length
- inference-time attention temperature scaling to enhance long-context reasoning and retrieval
🎉 Thrilled to share MLGym and MLGym-Bench, our new framework for AI Research Agents! 🚀 Developed during my Meta internship, MLGym provides a flexible environment for benchmarking and developing new agents for AI research tasks.
🔬 MLGym-Bench consists of 13 diverse AI research tasks, ranging from Language Modeling, Computer Vision, RL, Logic, and Game Theory.
We introduce MLGym & MLGym-Bench, a new environment and benchmark for AI research agents🤖, providing a standardized framework for evaluating LLMs on research tasks🧠🚀
📄 Full paper: https://t.co/R4huVjDKyB
Super excited to share 🧠MLGym 🦾 – the first Gym environment for AI Research Agents 🤖🔬
We introduce MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks.
The key contributions of our work are:
🕹️ Enables the exploration of different training algorithms for AI Research Agents such as RL
🛠️ Provides a flexible evaluation framework that can accommodate different artifacts such as models, algorithms, or predictions
🤖 Allows researchers to evaluate any model without the need to develop a custom agentic harness
🎯 Introduces 13 diverse open-ended AI Research tasks for evaluating AI Research Agents on a wide range of domains such as computer vision, natural language processing, reinforcement learning, game theory, and logical reasoning.
📈 Proposes a new evaluation metric for AI Research Agents
MLGym makes it easy to:
1) Add new tasks
2) Evaluate new models
3) Integrate new agents
Check out a video of the MLGym Agent to see how it performs the full pipeline of idea generation💡, implementation 👩💻, experimentation 👩🔬, and iteration 🔄 to improve on ML tasks.
Huge thanks to the exceptionally talented @deepaknathani11 who led this work and to all the other amazing collaborators who made this possible 🙏🫶🚀
Super excited to share 🧠MLGym 🦾 – the first Gym environment for AI Research Agents 🤖🔬
We introduce MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks.
The key contributions of our work are:
🕹️ Enables the exploration of different training algorithms for AI Research Agents such as RL
🛠️ Provides a flexible evaluation framework that can accommodate different artifacts such as models, algorithms, or predictions
🤖 Allows researchers to evaluate any model without the need to develop a custom agentic harness
🎯 Introduces 13 diverse open-ended AI Research tasks for evaluating AI Research Agents on a wide range of domains such as computer vision, natural language processing, reinforcement learning, game theory, and logical reasoning.
📈 Proposes a new evaluation metric for AI Research Agents
MLGym makes it easy to:
1) Add new tasks
2) Evaluate new models
3) Integrate new agents
Check out a video of the MLGym Agent to see how it performs the full pipeline of idea generation💡, implementation 👩💻, experimentation 👩🔬, and iteration 🔄 to improve on ML tasks.
Huge thanks to the exceptionally talented @deepaknathani11 who led this work and to all the other amazing collaborators who made this possible 🙏🫶🚀
Today, we’re introducing Jace, your AI Email Agent.
Jace uses your past responses, checks your calendar, and pulls in context from attachments or the web to draft replies in your voice and schedule your meetings.
Imagine an executive assistant that knows exactly what to reply, matches your writing style, and works tirelessly for you 24/7.
Before you even open an email, the perfect draft reply is ready and waiting for you to hit send.
That’s Jace.
Try Jace below 👇