Member of Technical Staff @MosiAI_Official | ex-Shanghai AI Lab Pretrain Team @intern_lm | Dev @OpenMMLab #MMRazor | Building open-source MLLM & ML systems
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!
🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
https://t.co/2YWSvJHhKA
Even in the vibe-coding era, the SGLang-Omni team continues to set the bar for rigorous systems engineering—solid profiling, careful optimization, and results that speak for themselves. Seriously impressive work. Huge kudos to the team! 🚀
The hardest part of maintaining an open-source project is serving them at scale while keep high code quality.
In SGLang-Omni, we spent a month refactoring, removing 2000+ non-test duplication. Here is what we changed, and what we learned. https://t.co/kaxYsYokmo
Count us in! SGLang community has co-signed the Open Weights letter @lmsysorg
Everything we ship goes out in the open, because world-class performance should be accessible to every builder.
Let's keep building in the open 🧡
The first Vera Rubin clusters are here!
Yesterday, @IneffableLabs took delivery of their Vera Rubin NVL72 cluster from @googlecloud@nvidia
The AI frontier jumps forward by yet another generation of hardware. Acceleration continues.
Banning Chinese models is the most un-American thing we could do.
Stop complaining and start competing.
People won't use Chinese models if American companies build something better.
When the SGLang Omni team profiled MOSS-TTS Local v1.5, we expected autoregressive decoding to dominate. Instead, the largest bottleneck was its transformer vocoder. I built a packed FlashAttention path that increased QPS by 48.64% and reduced mean latency by 32.71%.
@ShahzaibAli3029@MosiAI_Official@huggingface We welcome community contributions to bring our models to third-party frameworks like CrispASR. We’ll open a GitHub RFC soon, with our team actively supporting the effort. Stay tuned!
🤗 MOSS-Transcribe-Diarize-0.9B is now open source on @huggingface.
Built with an end-to-end audio-to-structured-transcript paradigm:
>0.9B open-source ASR model
>Apache license 2.0
>128k long-context transcription
>Up to ~90-min audio input
>Speaker labels + timestamps in one generation
>Multi-speaker diarization for meetings, interruptions, and overlapping voices
>Hotword biasing for names, terms, and domain-specific vocabulary
>~100 token/s on NVIDIA RTX 4090, RTF ~0.017
Thank you @sgl_project@vllm_project@Prince_Canuma@lllucas for day-0 support! 🚀
Github: https://t.co/0hgIjvhGv1
Huggingface: https://t.co/gQ1avC3w2e
API: https://t.co/znzfO4PHuu
Live demo: https://t.co/wAmEbbFna5
Technical Report:https://t.co/3yb0VKp7FS
HF Space: https://t.co/gQ1avC3w2e
AtomGit:https://t.co/uIwqnkMeGU
SGLang-Omni: https://t.co/TFlp4YzI7f
vLLM: https://t.co/j1VMVuQDcp
MLX-audio: https://t.co/ofDO9V6N7J
Discord:https://t.co/H3shzNLHOL
🚀 MOSS-TTS Local Transformer v1.5 is now available on Novita.
Build voice agents with:
🎙️ Zero-shot voice cloning
🌍 Speech synthesis in 30+ languages
🎧 Native 48 kHz stereo audio
⚡ Streaming TTS with ultra-low latency
Thanks to @MosiAI_Official for open-sourcing MOSS-TTS.
We’re glad to see the community building around MOSS-TTS Local Transformer v1.5.
A developer has created a ComfyUI custom node for MOSS-TTS, helping bring local text-to-speech generation into ComfyUI workflows.
Thanks to the contributor — 🌟https://t.co/Bjio3vtSMD