Elon Musk says Orbital AI Data Centers will be easier than Communication Satellites for SpaceX:
"The Starlink V3 communications satellites is an incredibly complex machine. The AI Data Center will be much simpler by comparison.
It’s really just solar power, plus radiator, some basic equipment, and the laser links would connect to the Starlink communications constellation, and then to the ground."
🎉Thrilled to announce EAGLE 3.1 - the next evolution of speculative decoding from @EagleCorp, developed by @hongyangzh, @dogacel0, and the EAGLE team in collaboration with vLLM @vllm_project and TorchSpec @lightseekorg!
💡EAGLE 3.1 introduces a new FC normalization + post-normalization hidden-state feedback architecture that significantly improves long-context robustness, acceptance length, and serving stability across real-world inference environments.
Shoutout to @NVIDIA who has been instrumental in the large-scale training, benchmarking, and inference validation of EAGLE 3.1 to help bring this next step in inference acceleration to production environments.
For EAGLE 3.1, the EAGLE team identified attention drift as a key bottleneck behind deeper-step acceptance-length degradation in speculative decoding.
✨What's new:
• Up to 2× longer acceptance length in long-context
• Stronger long-context + chat-template robustness
• More stable serving across diverse prompts or environments
• Native vLLM support
• TorchSpec training support
• Open-source Kimi K2.6 EAGLE 3.1 draft model
🔗 Blog: https://t.co/874soS61Sg
So we built fabric-lib: an RDMA point-to-point abstraction for LLM systems.
✅ Portable across NICs (NVIDIA ConnectX, AWS EFA)
✅ Delivers peak RDMA performance
✅ Designed for the messy realities of modern LLM workloads
Collectives are great for structured parallelism, but awkward for emerging LLM workloads:
1️⃣ Fixed membership — hard to add/remove replicas
2️⃣ Synchronized init — everyone blocks on setup
3️⃣ Shape uniformity — MoE padding = 50x waste
4️⃣ Strict ordering — forces extra buffering
I know, these days stars don't mean much, but it's still lovely to see such a nice round number. Thank you everyone for trusting your ML workloads to us, here's to the next 100K stars! (H/t Tianyu for pulling this screenshot)
Giving a talk on behalf of @vllm_project about open source at #MLSys 2026 tomorrow and will be around in Bellevue May 18-21. https://t.co/SEyl6Y5HbZ
The @inferact crew will be here too with a booth! Come say hi!🤗
@dwarkesh_sp@reinerpope this is awesome!! one nitpick: DeepSeek-V3 is ~18× sparse (671B total / 37B active). In expert terms it's top-8 routed + 1 shared
Codex grew programmatic policies with no neural nets: max score on Breakout, and SOTA-level scores on MuJoCo.
Maybe heuristics were not too weak. Maybe they were just too expensive to maintain. Maybe it's the next paradigm.
https://t.co/1ZaIneleuW
A lot of the stack we use today was built for an earlier paradigm. If we want new kinds of research, we’ll need new libraries, new abstractions, and new ways of running the lab.
I couldn’t be more excited about our founding team. We’re going to any% the frontier.
It’s been an exciting nine months training this model from scratch. I’m especially proud of the opportunity to rebuild the foundational infrastructure alongside the strongest infra team I’ve ever worked with. The systems we’ve built will serve as a solid foundation for many more models to come. Stay tuned!
Excited to share Muse Spark, the first model from whole team’s work in MSL! 🚀
It’s natively multimodal and agentic. I’ve been using it for my daily coding and research tasks. Still plenty of room to improve in agentic domains, but we’re moving with great velocity.
It’s a seriously good model! Check out the full breakdown and try it out in https://t.co/Fka0wdAswy
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵