muse code tips #1: ephemeral worktrees
muse -w creates an isolated worktree under .muse/worktrees/
auto-removed when safe, retained with a reason if not
1/ we just publicly released Muse Spark 1.3 max!
we see significantly stronger coding and agentic performance on muse spark 1.3 max, so would strongly recommend trying it out even if you've already tried muse spark 1.3 high or muse spark 1.3 xhigh.
Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities
Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62)
Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score
Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release!
Key Takeaways:
➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh)
➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available
➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh)
➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh)
Other model details (xhigh variant):
➤ Context window: 1M tokens, unchanged from Muse Spark 1.2
➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M
➤ Input modalities: text, image, video
➤ Availability: Meta's first-party API and Muse Code
One-shotting games might be more fun than playing them 😆
We put a lot of work into visual reasoning and coding for Muse Spark 1.3 — excited to see what games people build with it.
Any RTS fans here? 👀 Here's one we built with it:
🎮 https://t.co/tc4hx7J96S
More in our report: https://t.co/ZRWzKDPDjJ
2/It's a step up on reasoning and coding. compared to Muse Spark 1.2, Muse Spark 1.3 wastes fewer turns, uses ~20% fewer tool calls, ~25% fewer tokens, and holds onto requirements well during long-horizon tasks.
Very excited to release Muse Spark 1.3! The coding and long horizon performance got a huge jump! 🚀
It's still incredibly cheap and also much better at visual coding 🥳
Try it out!
https://t.co/7XOQ8lSw1i
Muse Code is out of beta and now built to handle bigger, more complex engineering tasks. Developers can get started with one command today:
curl -fsSL https://t.co/0RApZrEJMv | bash
The bridge between agentic tool use, multimodal understanding, and real-time visual coding opens up incredible possibilities for long-horizon tasks.
So proud of our team's work on the game featured in the post! Play the game here: https://t.co/wcdNUY2pkj
1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.
Muse Spark 1.2 is a frontier model for game development. It matches GPT-5.6-Sol, tying for third place.
Earlier this year, Meta didn’t even have a model capable of making this leaderboard. A huge leap for the team—congrats on an amazing model!
To understand whether we're making genuine progress on reasoning, we entered our AI models in five international STEM Olympiad competitions this year.
The results:
🏅 Asian Physics Olympiad (APhO): Perfect score on the theory exam — gold medal
🏅 International Physics Olympiad (IPhO): Perfect score on the theory exam — gold medal
🥇 International Mathematical Olympiad (IMO): Gold medal, top 4% of human participants
🥇 International Chemistry Olympiad (IChO): Gold-medal level performance
🥇 Romanian Masters of Mathematics (RMM): Gold-medal level performance
Three of these (APhO, IPhO, IMO) were live competitions and our solutions were submitted under real competition conditions and graded by the official judges using the same marking criteria applied to student contestants.
A few things about the approach:
• Models were internally trained versions from the Muse Spark family
• Zero tool use: no search, no code interpreter, no calculator
• Multi-agent orchestration with parallel reasoning
We are excited about where this reasoning capability goes next; frontier research level across scientific domains and personal superintelligence.
Super grateful to the organizing committees of APhO, IPhO, and IMO for supporting our live participation. We have deep respect for the contestants and organizers behind these competitions. 🙏
And proud of the MSL team that pulled this together!
The Muse Spark 1.2 API pricing is quite crazy, especially as a daily driver for most coding tasks. It's up to 250x cheaper than Fable and 150x cheaper than GPT-5.6 Sol!
Price per 1M tokens:
Cached input - $0.002
Input - $0.10
Output - $0.20
Try it out!
https://t.co/8QEHL7rjIr
AI-generated games and real-time environment creation are moving way faster than expected! It is fun to play with Muse Spark 1.2 + Muse Code. Please try it out 🎮
To play: https://t.co/EvcZaArdSC
Check out our blog post: https://t.co/J92B0Biv6b
Muse Code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task.
When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees. Your working copy is never touched. In testing we had it build six features for a game simultaneously with no collisions.
Muse Code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task.
When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees. Your working copy is never touched. In testing we had it build six features for a game simultaneously with no collisions.