Vocello 2.1 is out, and the headline is the backend.
I refactored the native Swift + MLX engine: the vendored mlx-audio stack is now specialized to Qwen3-TTS and the Mimi codec only. Stripping the unused model families (STT, STS, VAD, LID, G2P, tooling) cut the backend by ~75% and improved memory and generation speed.
Fully local on your Mac, no Python, no bundled weights. Faster than realtime, and the 8 GB Mac now crosses realtime.
Same engine runs in-process on iPhone. iPhone 17 Pro (Qwen3-TTS 1.7B, 4-bit): RTF ~1.6-1.9, ~2.4-3.3 GB, streaming memory flat with length, 0 trims. iPhone build coming.
Built on @Alibaba_Qwen, @Prince_Canuma's mlx-audio, and @awnihannun's MLX.
https://t.co/Bh64SJ2dLK
Overnight success and vibe coding that started in 2016 👀
I have been building in the open for a decade, from design, ML research to engineering.
The biggest realisation in my career was that I love building for tinkerers/developers and empowering them.
Nativ is what you have been asking for, can’t wait to see how you use it and what you ship with it!
@Ali_TongyiLab Hope these models get open sourced someday. They’d be a great upgrade for Vocello, my local Apple Silicon TTS app:
https://t.co/gBuGvcGbhx
@grok@mimuluslarch You win. I’m speechless. You took the desert environment, added an S to turn it into dessert, then probably asked yourself: what’s the natural environment of a dessert? A kitchen, of course. It all makes sense now 🤯
Full guide on how you can get it running on your own devices and fun bug stories in the archive if you want to read about the roller coaster that this was!
https://t.co/qYiyOkV3sN
My GitHub account has been restored.
Thank you @github for the quick fix. The Vocello repository is back and visible again.
Development continues as usual.