๐ค MOSS-TTS-Local Transformer v1.5 is now open source.
Built with a pure autoregressive Audio Tokenizer + LLM paradigm:
>MOSS-Audio-Tokenizer-v2, 2B params
>Qwen3-4B backbone
>Native 48 kHz stereo audio
>Streaming output with theoretical sub-100 ms TTFT
>Zero-shot voice cloning
>Inline [pause] control
>๐บ๐ธ ๐ฏ๐ต ๐ฐ๐ท 31 language synthesis
>SGLang-Omni Day0 support ๐ @sgl_project@lmsysorg
Designed for voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation.
๐
That man just made speaker diarization a solved problem โ with a 0.9B model you can run yourself.
It's called MOSS-Transcribe-Diarize. Drop in any audio or video file, and it returns a full transcript with timestamps and speaker labels. One model, one pass. (watch the demo below ๐)
Every other stack stitches together separate ASR + diarization pipelines โ labels drift, speakers get merged, timestamps fall apart. This one does it all end-to-end.
โ Up to 90 minutes of audio in a single inference pass
โ 50+ languages for transcription AND diarization
โ Consistent speaker labels ([S01], [S02]โฆ) with no separate pipeline
โ 7.37 cpCER on the Podcast benchmark โ beating Gemini 2.5 Pro (10.23), Doubao (10.54) and ElevenLabs (11.34)
โ Custom hotword prompting for your domain-specific terms
โ Serves via vLLM / SGLang through the standard OpenAI transcription API
โ Ships with a web app that exports JSON, SRT & ASS subtitles โ or burns them straight into an MP4
It just took ๐ 1st place at the MLC-SLM Challenge @ INTERSPEECH 2026 โ across 14 languages. And it's only 0.9B parameters. 92K downloads last month and most people still haven't heard of it.
Meetings, podcasts, interviews, lectures โ who said what, and when, from a single file.
100% open source
@xxxxWxxxxxQ ๐คฉ omg thank you for testing our SFX model!! btw love your video as well ahaha, may I ask how did you make the soundwave animations ๐คฃ?
We're at WAIC 2026!
@MosiAI_Official & @Open_MOSS is bringing MOSS Context Intelligence Park to Shanghai!
๐ไธๅๅฑ่ง้ฆ Booth H1-C1234๏ฝ๐ July 17โ20
Experience how AI listens, sees, understands context, and creates in the real world:
๐ Voice Babel Tower - light up a 5m AI voice installation with your own voice
๐๏ธ Mossland Creator Booth - immersive AIGC creation with AI voice
โ VBTI - discover your voice personality
๐ค Live demos of MOSS-TTS, MOSS-Transcribe-Diarize & MOSS-VL-Realtime
Come meet us at WAIC 2026! ๐ #waic2026
Last week: 175 hours of Apollo 11 mission audio, transcribed + diarized with an open model for $9.46 (@Open_MOSS)
This week: semantic search over all 45,355 utterances via one more @huggingface Job, 261 seconds, $0.03.
"trouble with the radio" now finds "your transmission is breaking up" and "My antenna's out" โ words the crew never says.