We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model, trained this way.
Today we’re introducing MIRA, a new multiplayer world model, built with @gen_intuition, in collaboration with Epic Games.
We release an in-depth technical report, dataset, as well as an online demo that you can try right now (link below).
Pocket TTS goes multilingual!
Now you can use our 100M-parameter models to generate speech in six languages, fast enough to run real-time without a GPU. We also improved the quality of the English model while keeping the same size. And all of this is open-source.
We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required.
Open-source, lightweight, and incredibly fast. 🧵👇
Kyutai Speech-To-Text is now open-source! It’s streaming, supports batched inference, and runs blazingly fast: perfect for interactive applications.
Check out the details here: https://t.co/bQMP56XaKC