Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.
Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come.
But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility.
This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems.
Together, we are building the foundation of the AI economy.
Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://t.co/ugYWQ1MyRi
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today.
Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once.
Compared with MiMo-V2.6's Hybrid SWA architecture:
• 5.02× lower prefill FLOPs at 1M tokens
• 4.5× smaller KV cache at 1M tokens
• Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL
Why build a new architecture?
Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time.
HySparse2 tackles all three with two levels of KV sharing:
• KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states.
• KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices.
Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes.
Paper: https://t.co/REeEdd5pL7
Cannot understand why CDG Engie and CDG Energy cannot get their app right and their charging points properly maintained. Today the app failed with "Internal Server Error" when using the Add Payment Card!!! We need stability and prompt support CDG. Please get your act together!
AI for science is one of the greatest positive forces we have, and I cannot think of anything more human than to understand nature and to use the power to create new technologies that improve our lives, civilization and allow us to reach beyond. https://t.co/Cc6ZSRBJgz
Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀
- 3 custom chips: Xring O3, O100, D100
- 200 TOPS NPU
- 1.22 TB/s AI memory bandwidth
- Up to 160GB unified memory
- 150W sustained power
- 120B models running locally
Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥
Once it hit's the marked, going to sell like hot cakes.
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.
Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.
Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows.
Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs.
In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license.
🧵👇