@YottaLabs partnered with @radixark to deliver the first @sgl_project inference on @awscloud Trainium. This isn't just a bold step toward embracing a multi-silicon, multi-cloud future—it's a strategic move to significantly reduce token costs for end users.
Check out the blog and open-source code: https://t.co/PGD0UGfAl1
Btw, you can also launch a Trainium instance on Yotta Console with 1 click: https://t.co/iN9z3c2dPq
Multi-silicon is inevitable, our vision is to build the best ecosystem for developing AI native applications without worrying about the underlying infrastructure complexities.
Our CEO Da Li took the stage at Seattle Tech Week today.
His case: multi-silicon already happened, and the software layer is the new high ground in AI infrastructure.
The line that stuck with the room: the question is no longer "which chip should we buy." It's "does our software let us choose freely?"
This is what we're building for.
Tonight in Bellevue: the Seattle World Models Carnival, co-hosted with Physion Labs.
World models, video generation, AI infra. Dinner and demos with the people building this space.
It's a small dinner and spots are limited. Request an invite: https://t.co/ZgNvig5qzr
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
🚀Hy3 is here.
295B MoE. Best in its size class. Rivals trillion-scale flagships.
Reliable and affordable for most agentic usecases.
Apache 2.0. Friendly for commercial use.
FREE API for 2 weeks → https://t.co/EyURKwTdgi
🤗 https://t.co/twqJpqb2SL
📖 https://t.co/4uEkIU1cW4
DSpark marks a new era where spec decoding no longer slows down under high concurrency.
We took time to carefully optimize variable-length verification, keeping it fast while maintaining high verification quality.
Run Qwen3 and DeepSeek-V4 with DSpark in SGLang now, let's goooooooo 🚀
I completely stopped coding, and I stopped clicking the fancy UI, I just ask my coding agent to do everything for me, including launching a GPU pod and run more AI. Btw, GPU doesn’t equal NVIDIA on Yotta, it is just a computing silicon.