Congratulations to the Ling team @AntLingAGI on the release of Ling-3.0-flash! 🎉 A 124B MoE model activating just 5.1B parameters per token—with hybrid-linear attention and native 256K context—is an exciting contribution to production-scale agent systems.
We also applaud the team’s announce-first, open-source-next release approach. Separating the announcement from the weight release lets the model team freeze the final checkpoint, configuration, tokenizer, and serving semantics, while giving open-source inference projects a stable window for correctness testing, performance tuning, Docker builds, and recipe validation. The community also gets a clear timeline instead of an ambiguous “coming soon.”
This is not a step back from day-0 support—it is a more sustainable way to deliver day-0 support for the artifacts users will actually run.
vLLM’s open-source support for Ling-3.0-flash is coming soon and will be available when the model weights are open-sourced. We hope more model vendors adopt this release pattern in the future!
Congrats again—we can’t wait to see what the community builds! 🚀
AI ENGINEER RETURNS TO MIAMI!
Bringing 500 AI engineers, CTOs, and VPs of AI back to Miami to connect, learn, and build what's next in AI engineering.
Early Bird Ticket Sales - OPEN
Sponsorship Inquiries - OPEN
Call for Papers - OPENING NOV
↓↓↓
Just found something awesome! If you self host or deploy llms in production this is beyond incredible. I am sure most of the big labs either use this or some custom implementation. Save those prefix cache hits for reuse!
https://t.co/1bfpoAPG0Y
@ivanfioravanti glm 5.2 nf3 variant running on 4x6000s Interesting to compare outputs I assume yours was from api glm 5.2 ? this is from https://t.co/iIsicwqEwT
@0xSero@digitalix the hybrid is the best choice, recently updated weights are so close to the original performance. To be fair I did not load the 2bit variation.
We're generating 6 second video in ~6 seconds, 1080p, 144 frames with LTX 2.3 Fast. That's 4x faster than GPUs.
ICYMI other new performance benchmarks incl:
- DeepSeek-R1-0528 671B: 400+ t/s/u
- Kimi K2.6: 900 t/s/u, 3x faster than GPUs [soon!]
Try the latest on https://t.co/lb0XovgEHI
OpenCode Manager lets you work with OpenCode across all your projects — from terminal or phone.
Use ocm to push/pull repos, attach your local TUI to Manager, or move a live session to mobile.
Then manage Git, browse files, control assistants, use voice input/output, and keep going anywhere.
https://t.co/1HSGvFPBph @opencodeecho@opencode