Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows.
Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs.
In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license.
🧵👇
I believe everyone should have access to superintelligence, and I wrote a long piece about Meta's philosophy and values for building a positive future for everyone. https://t.co/2ZoNZXZ39T
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Open source movement made software something everyone could build on. IMO, AI is heading the same way, frontier models will keep pushing the ceiling and open weights will spread what’s underneath them.
Jensen is right that we need both and real win is AI showing up everywhere.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
The way to think about “open” in software is being a low-cost producer (vs a high margin one). When a company (or 2) achieves record valuations in record time, that rightfully attracts competition (as it should). We need to let the free market work.
https://t.co/OJ9XPcKmX4
The new Kimi K3 leverages LatentMoE, a technique developed by @nvidia in January
What is LatentMoE? In LatentMoE, tokens are projected from the model hidden dimension 𝑑 into a smaller latent dimension ℓ for expert routing and computation, which reduces routed parameter loads and all-to-all traffic by a factor of 𝑑/ℓ.
Learn more about it here: https://t.co/n9vr4fhHLw
Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 👇️ 1/8🧵
Wow! Kudos to the moonshot ai team for Kimi K3. It’s like a second deepseek moment 🍿 for the AI era.
No doubt, US frontier models will catch up, 🇺🇸 is still the best place to responsibly push the frontier.
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
I'm biased, but I wish I could give this 100 likes.
The idea that really hit home for me:
If we had the ultimate AI, would it go and reinvent ServiceNow or SAP or other software tools? Or would it just use the proven tools out there because that's the most efficient way for it to achieve its goals? It would use the proven tools.
I also like that he framed "tool use" as one of the biggest advancements in AI. He's totally right. Once we started giving AI access to tools, it became much more effective.
<rant>
If I have an AI agent working in my GTM team, and I ask it how many customers we signed up last month, I don't want it to reason over a bunch of the Internet, read reddit comments and scan internal emails and make its best guess at the answer. I want it to just ask the CRM (like @HubSpot) that was specifically built for that purpose and give me the definitive answer.
If I had a 200+ IQ human join my team and I asked them to summarize sales in Europe last quarter, I want them to access the system of record, not start vibe coding its own CRM.
</rant>
@AirIndiaX screwed up rescheduling my ticket from Ahmedabad-BLR purchased through AMEX today, asked me to buy a new ticket due to their technical blunder. I ended up buying a new ticket to make it to my destination. Get your systems right, your customer experience is horrible.
Congrats to the xAI team on Grok 4.1. Topped LMArena at launch, 65% fewer hallucinations, strong creative benchmarks. Solid execution.
@elonmusk never fails to amaze! 🎉
Introducing Gemini 3 ✨
It’s the best model in the world for multimodal understanding, and our most powerful agentic + vibe coding model yet. Gemini 3 can bring any idea to life, quickly grasping context and intent so you can get what you need with less prompting.
Find Gemini 3 Pro rolling out today in the @Geminiapp and AI Mode in Search. For developers, build with it now in @GoogleAIStudio and Vertex AI.
Excited for you to try it!