Meet Nex-N2.5 family, our latest open source agentic models.
Mini (35B) and Pro (397B) bring multimodal understanding and computer use. Max (1.6T) is a text-only MoE model built for complex reasoning, coding, and agent workflows.
Strong benchmark results in workflow automation and computer use:
⚡ Max scores 50.2 on AutomationBench v1.0.6—just 0.1 points behind Claude Opus 5
⚡ Pro scores 56.4 on OSWorld-2, ahead of Qwen3.8-Max at 46.7
With continuous action, visual feedback, and self-correction, NEX-N2.5 can work across Blender and CAD tools, handle expense reimbursements, and even play PC games—turning visual understanding into sustained, real-world execution.
🔗@huggingface
https://t.co/EiXIHqMHh7 https://t.co/chy2SqwUu9 https://t.co/dasDLotZbK
🔗@modelscope
https://t.co/pQQzrXbPuN https://t.co/PSRiuE9Dmr https://t.co/p5QELC4rQd
🔗 Github https://t.co/MtkGobY8Nn
🔗Website https://t.co/7oLSfyOCxB
🤗 MOSS-Transcribe-Diarize-0.9B is now open source on @huggingface.
Built with an end-to-end audio-to-structured-transcript paradigm:
>0.9B open-source ASR model
>Apache license 2.0
>128k long-context transcription
>Up to ~90-min audio input
>Speaker labels + timestamps in one generation
>Multi-speaker diarization for meetings, interruptions, and overlapping voices
>Hotword biasing for names, terms, and domain-specific vocabulary
>~100 token/s on NVIDIA RTX 4090, RTF ~0.017
Thank you @sgl_project@vllm_project@Prince_Canuma@lllucas for day-0 support! 🚀
Github: https://t.co/0hgIjvhGv1
Huggingface: https://t.co/gQ1avC3w2e
API: https://t.co/znzfO4PHuu
Live demo: https://t.co/wAmEbbFna5
Technical Report:https://t.co/3yb0VKp7FS
HF Space: https://t.co/gQ1avC3w2e
AtomGit:https://t.co/uIwqnkMeGU
SGLang-Omni: https://t.co/TFlp4YzI7f
vLLM: https://t.co/j1VMVuQDcp
MLX-audio: https://t.co/ofDO9V6N7J
Discord:https://t.co/H3shzNLHOL
🤗 MOSS-TTS-Local Transformer v1.5 is now open source.
Built with a pure autoregressive Audio Tokenizer + LLM paradigm:
>MOSS-Audio-Tokenizer-v2, 2B params
>Qwen3-4B backbone
>Native 48 kHz stereo audio
>Streaming output with theoretical sub-100 ms TTFT
>Zero-shot voice cloning
>Inline [pause] control
>🇺🇸 🇯🇵 🇰🇷 31 language synthesis
>SGLang-Omni Day0 support 🎉 @sgl_project@lmsysorg
Designed for voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation.
👇
MOSS-TTS-v1.5 just reached #1 on Hugging Face Trending for Text-to-Speech, with 20.6K downloads.
A multilingual, controllable TTS model with stable voice cloning, long-form generation, and precise pause control.
MOSS-TTS-v1.5 is now officially supported by vLLM-Omni and SGLang-Omni.
Built by OpenMOSS-Team.
Try it:
GitHub: https://t.co/mSlALD6Fzy
Hugging Face: https://t.co/qTv7xu1MZ5
ModelScope: https://t.co/NzAXgAzagL