We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model, trained this way.
We're releasing Voice of Reason, a speech-native model that does math out loud. Give it a spoken problem without transcription nor text LLM in the loop, and it reasons and answers in speech. GSM8K goes from 27.3% for GLM-4-Voice to 77.1%. Link in 🧵
If you're looking for a weekend project, how about training your own text-to-speech model from scratch on your own GPU, and then running it on any device's CPU?
We just open-sourced the entire Pocket TTS training stack: data pipeline, recipes, and evals.
It learns pretty damn fast:
~15k steps: babbling starts turning into words
~50k steps: it reads anything you type (WER under 1%)
~200k steps: the voice stops sounding synthetic
On a beefy consumer GPU, that's a week of training. On eight H100s: 10-20 hours. A TTS training run will cost you less than $200 if you rent your hardware, and an order of magnitude less if you just pay for power.
Some things we'd love to see people try:
- Train it in your own language (a few hundred hours of speech gets you surprisingly far).
- Add new features to Pocket TTS (Emotion tags? Make it sing?).
- Beat us at our own game: make it faster and smaller.
Show us what you build! We'll highlight the best models and new languages for the whole community to enjoy. Pocket TTS has already found many use cases, from reading for people with visual impairments to making NPCs in video games talk, and we're sure there's much more to do with it!
Here's an example of a Czech Pocket TTS. Try just asking your favorite agent to find data and apply the method, and you can have your own.
Get started: https://t.co/3EH3sbKNRU
We're releasing MuScriptor, the best open model for multi-instrument transcription to date, created in collaboration with @MireloAI.
Give it a recording in any genre: pop, classical, metal, jazz, whatever, and it transcribes the individual instruments into MIDI. Link in 🧵
Heading to ICML 2026 in Seoul next week with @romfbr31 to present Hibiki-Zero🇫🇷🇬🇧🇵🇹🇪🇸🇩🇪[https://t.co/D7gadZ36Ib], Kyutai's latest real-time speech translation model.
I'll be giving an oral presentation on July 8 at 10:30 AM KST. Feel free to join if you'd like to learn more!💬
Today we launch stt-translate and s2s-translate: real-time speech-to-text and speech-to-speech translation. They compete with gemini-3.5-live-translate and gpt-realtime-translate on latency and quality, while allowing you to speak in any voice from our catalog or one you clone. Try them for free today on https://t.co/HLVvh94Kok
Three of our papers got accepted at ICML and one at CVPR this year 🎉 We will have researchers on-site for both conferences, so come talk to us if you want to learn more about Kyutai!
👁️ MoshiVis (CVPR’26) → Vision Speech Models: A data- and training- efficient pipeline for omni models built on top of Moshi
🧠 MoshiRAG (ICML’26) → Making speech-to-speech models smarter with the power of RAG and minimal latency
🗣️Hibiki-Zero (ICML’26) → Streaming speech-to-speech translation without aligned data leveraging GRPO
⌛ Kairos (ICML’26) → Recency bias is real, even for LLMs. More details in a future post!
#ICML2026 #CVPR2026
Speech-native models like Moshi sound great and answer fast, but aren’t as smart as text LLMs. In our new paper, MoshiRAG, we show how Moshi can ask for advice from a text LLM or a knowledge base. The tricky part is how to do this in real time without adding latency. 🧵
Gradium builds models, not orchestration or voice agents. But to really evaluate the conversational experience around our models, you need to see them inside an actual agent.
That’s why we built Gradbot internally, to spin up a POC in minutes before a sales call.
Now we’re open-sourcing it for anyone to experiment with and have fun.
We're releasing OVIE, a novel view generation model trained entirely on single images. No multi-view datasets needed.
Given a single image, it generates novel views of any scene in real time, running orders of magnitude faster than competing approaches.
Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications
My conversation with @neilzegh, top AI researcher in the field (@GoogleDeepMind, @Meta, @kyutai_labs) and now CEO of @GradiumAI
This is a reference episode on all things voice AI 🔥
00:00 Intro
01:21 Voice AI’s big moment, and why we’re still early
03:34 Why voice lagged behind text/image/video
06:06 The convergence era: transformers for every modality
07:40 Beyond Her: always-on assistants, wake words, voice-first devices
11:01 Voice vs text: where voice fits (even for coding)
12:56 Neil’s origin story: from finance to machine learning, with help from @ylecun and @soumithchintala
18:35 Neural codecs (SoundStream): compression as the unlock
22:30 Kyutai: open research, small elite teams, moving fast 31:32
Why big labs haven’t “won” voice AI4
34:01 On-device voice: where it works, why compact models matter
46:37 The last mile: real-world robustness, pronunciation, uptime
41:35 Benchmarking voice: why metrics fail, how they actually test
47:03 Cascades vs speech-to-speech: trade-offs + what’s next
54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos
1:00:50 New languages + dialects: what transfers, what doesn’t
1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game
1:07:27 What data do you need to train voice models
1:09:02 Deepfakes + privacy: why watermarking isn’t a solution
1:12:30 Voice + vision: multimodality, screen awareness, video+audio
1:14:43 Voice cloning vs voice design: where the market goes
1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next
🌐 @tom_labiausse and @neilzegh just released Hibiki-Zero, a live translation model with a few seconds latency, and trained without any aligned audio data thanks to reinforcement learning.
Code, paper and checkpoints are out 👇
https://t.co/ErbJTKQL0d
We're releasing Hibiki-Zero, a new real-time and multilingual speech translation model that can translate 🇫🇷French, 🇪🇸Spanish, 🇵🇹Portuguese and 🇩🇪German to English: accurate, low-latency, high audio quality, with voice transfer. And best of all: open-source.
We hit 1k GitHub stars in 3 days with Pocket TTS, our 100M-parameter TTS with voice cloning that runs on CPU! According to internal estimates, we are on track to reach 182k stars by the end of the year.
Kyutai keeps shipping state-of-the-art open models pushing the frontier of voice research. This time: the first high-fidelity TTS that runs on CPU. Science and engineering are in Gradium’s DNA, can’t wait to see what the community builds with ultra-compact, on-device voice models.
Introducing Pocket-TTS, the first ever TTS model that runs in real-time on CPU (!) with high-fidelity voice cloning. Built on Continuous Audio Language, the newest wave of audio generative models from Kyutai. Under the guidance of Gradium's CSO @honualx who keeps leading audio research at Kyutai, the lab keeps pushing the frontier of research along Gradium's products.
We’re excited to introduce Pocket TTS: a 100M-parameter text-to-speech model with high-quality voice cloning that runs on your laptop—no GPU required.
Open-source, lightweight, and incredibly fast. 🧵👇
🏠 Introducing CASA: a new way to input visual information into LLMs. The current default to do that is by inserting image tokens into the text stream, but when using many images in long conversations, this floods the context window and is thus impractical for streaming inputs.🧵