Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier.
First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks.
- It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities.
- It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks.
- And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end.
Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing.
Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI.
- Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost.
All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat.
Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost.
Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare.
Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: https://t.co/v65eop5Ixq
Easily the most inspiring and humbling technical effort I've been a part of, all thanks to this team 🫶
The deliverable wasn't one great model (nor 7), but the muscle to build from scratch and the plumbing to keep them coming.
Hope you enjoy the read! https://t.co/Enhz88NMg2
Today, we're cutting the ribbon on our brand-new Standard Bots factory in Glen Cove, NY.
Come by for the live demos, robots in action, and meet the team building what’s next in automation.
📍 Factory tours all day
🎉 Talks + tech showcase starting at 10AM
👋 Walk-ins welcome
We just launched our dev blog @solver_ai! Check out my first two posts touching on our product vision:
- "The case for more rope (and context)": a false dichotomy btw scope & autonomy https://t.co/nqvP2XMAE1
- "Long live the IDE": how SWE will evolve https://t.co/OBsRo6hjnH
We're on @ProductHunt today! 🚀 Solver is your autonomous coding assistant that completes development tasks while you focus elsewhere. New features: image inputs & web research capabilities. We'd love your upvote! https://t.co/bZY2DLuEZ0 #ProductHunt#AIcoding
Let's gooo! Internally we can no longer imagine doing our jobs without it.
If you want to experience a system pushing the async + autonomous quadrant and freeing you up to do higher-multiplier work, come give us @solver_ai a try 🚀
Solver is now GA! 🎉
Our AI coding agent helps developers offload tasks completely. Set it, forget it, come back to completed work.
New: credit system, mobile experience, memories, Slack integration & projects feature.
Free tier available now at: https://t.co/MXGGXxtyrM
@raspberryjamchi@pirroh If you're interested in this space, consider giving us @solver_ai a try! We're pushing the limits on the async+autonomous quadrant of SWE agents. Would love to hear your thoughts 👉 https://t.co/8ErmUJaCYk
@pumatheuma@ameensol@drakefjustin based rollups would inherit the L1's censorship resistance and liveness, no? which has implications for liveness-related security (e.g. liquidation of bad debt)
Aaaand this is what we've been doing with my life. Couldn't be more proud of this crack team 🥹
Grab a spoon folks -- the proof is in the pudding (and all 9 of these showcase videos).
LFG 🚀
@LoganJastremski If ethereum had followed the original sharding roadmap, there'd be (I hope) no question that txs/tvl on the shards count towards ethereum metrics, right? Then an argument that (proper) L2s shouldn't be included is purely political, rather than technical.
@blairlmarshall Yea I don't dismiss the concern, just calling out that the economics of shared resources are reasonable, albeit inconvenient. Though there's an argument that as the ecosystem matures, we can expect higher but less variable fees on L1 blockspace, no? i.e. law of large #'s
@blairlmarshall Thought-provoking, because I don't think I agree. It *does* make sense that unrelated accounts compete for shared resources, unless one of the following is true:
1) we can scale shared resources arbitrarily at fixed or decreasing marginal cost
2) we choose to rate limit accounts
@penumbrazone Congrats on the milestone 🎉 But like, why use a new top-level domain?? Breaks chain-of-trust and normalizes a pattern that *will* get exploited by scammers
@PabloSabbatella @Almavieja369 @safe Estoy de acuerdo. Quiero creer que el UX de Safe (o derivados como Coinbase Smart Wallet) van a seguir mejorando y simplificando. Tmb es cierto que smart wallets introducen otro punto de falla y complejidad, pero cada día que pasa Safe esta mas battle-tested
@PabloSabbatella @Almavieja369 Un hardware wallet te debería dar tanto miedo como un EOA hot wallet. Los vectores de phishing y torpeza con seed phrases son mucho más prevalentes que software bugs. En mí opinión una multi-sig (eg @safe), idealmente con un HW signer, es lejos la mejor opción default 🤷
@BrantlyMillegan Then I'd go with an L2 to start. Probably @base, where I see a heterogeneous social community maturing much more rapidly because it's an accessible L2 (both from UX and cultural standpoints)