Many generative models work in an autoencoder's latent space; They train faster and generate better when this space is well organized.
One way to organize it is distilling a pretrained model into it. But how can it work when the two spaces have different structures?
We release a technical blog on training PocketTTS, our 100M parameter on-device TTS, through drifting, a recent one-step generative objective from Deng et al. Less than 1% WER, high quality and voice cloning. To the best of our knowledge, it's the first speech model, and the first autoregressive model, trained this way.
Meta is making a tamagotchi for its AI? Cute, I've had a few on my desktop for a few weeks 🐣
Tamaclaudchi is a (cute) Claude session manager to see and monitor all your Claude sessions at once
Made in collaboration with @romfbr31 and more features to come with @RamaAdrien
We're releasing Voice of Reason, a speech-native model that does math out loud. Give it a spoken problem without transcription nor text LLM in the loop, and it reasons and answers in speech. GSM8K goes from 27.3% for GLM-4-Voice to 77.1%. Link in 🧵
ECCV workshops scheduler webapp: https://t.co/jefmBGSiwz
Kyutai at ECCV 2026:
🏠 CASA (https://t.co/2W86xiDZ8q) @royaleerieme@MoritzBoehle J. Marrie
👁🗨 OVIE (https://t.co/ixQcjn1eZq) @RamaAdrien
🎲 The FID Lottery (https://t.co/ZJZzwCxMWw) @nico_dufour
🏎️ MIRA (https://t.co/QM2O08uey0) @RamaAdrien@vvolhejn
🗣️ AURA (https://t.co/DcL3LcqjLb, more details to come) N. Cohen
Several of us are at ECCV this week 🇸🇪. Here is a visual recap of where to find us👇
And if you haven't planned your workshop days yet, @nico_dufour made a webapp over the weekend to get a comprehensive view of all speakers+events across workshops. Link below + more details soon
After 2 weeks of development, I finished optimizing the Mimi codec model, and it runs 15.4x faster for FP32.
Many omni and TTS models use the Mimi codec model. So I will fine-tune it to improve the WER and DNSMOS quality of the Mimi codec model and release it as open source.
The biggest obstacle for open source work is GPU. If you want to sponsor, you can send me a message.
I'm slowly accepting never being able to read all of my open arXiv tabs
So I'm building PeperNoten 🍪 paste an arXiv link, it bakes the paper into a first read-through Obsidian note. tl;dr, method, results, figures, bibtex.
Named after the Dutch cookie, for extra cuteness
We're releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training.
- 1.3B video URLs from CommonCrawl
- 80M downloaded videos
- 10M video hours
- 55M captioned clips
- 300M frame-caption pairs
🌐: https://t.co/hV2b92QSKd
If you're looking for a weekend project, how about training your own text-to-speech model from scratch on your own GPU, and then running it on any device's CPU?
We just open-sourced the entire Pocket TTS training stack: data pipeline, recipes, and evals.
It learns pretty damn fast:
~15k steps: babbling starts turning into words
~50k steps: it reads anything you type (WER under 1%)
~200k steps: the voice stops sounding synthetic
On a beefy consumer GPU, that's a week of training. On eight H100s: 10-20 hours. A TTS training run will cost you less than $200 if you rent your hardware, and an order of magnitude less if you just pay for power.
Some things we'd love to see people try:
- Train it in your own language (a few hundred hours of speech gets you surprisingly far).
- Add new features to Pocket TTS (Emotion tags? Make it sing?).
- Beat us at our own game: make it faster and smaller.
Show us what you build! We'll highlight the best models and new languages for the whole community to enjoy. Pocket TTS has already found many use cases, from reading for people with visual impairments to making NPCs in video games talk, and we're sure there's much more to do with it!
Here's an example of a Czech Pocket TTS. Try just asking your favorite agent to find data and apply the method, and you can have your own.
Get started: https://t.co/3EH3sbKNRU
Our audio-to-MIDI model, MuScriptor, now also detects tempo! You can directly drag-and-drop the MIDI into a DAW and it just works, without you having to tempo-match manually. Have fun!
Good time to remind everyone that @kyutai_labs has been one of the most prolific labs for open audio models and papers in its 2.5 years of existence. Detailed research papers that not only boast about metrics but also explain methods extensively have become quite rare. By submitting our papers to peer-reviewed venues, we make sure we meet the replicability bar. Even after creating @GradiumAI, I’m still committed to this mission through the work of my talented students!
Phonon (100M params) already led on English. Now it has a lower word error rate than both NVIDIA's Magpie (357M) and Neuphonic's NeuTTS Nano (229M) in French, German, and Spanish too, with high quality voice cloning, at a fraction of the size. Full multilingual benchmarks: https://t.co/5gNmlIgJzH
Get it now: https://t.co/8Tu3xwHj8d
We're releasing MuScriptor, the best open model for multi-instrument transcription to date, created in collaboration with @MireloAI.
Give it a recording in any genre: pop, classical, metal, jazz, whatever, and it transcribes the individual instruments into MIDI. Link in 🧵
Big day! I'm proud of how far we managed to take this project (you should've seen the baseline models). Try out the world model, it's really something.
My team open sourced a set of agent skills to run an entire @kaggle competition workflow from plain language. We ship an easy to install plugin for your favorite coding agent.
👉 Try it now: https://t.co/ayVnzYycLY
Heading to ICML 2026 in Seoul next week with @romfbr31 to present Hibiki-Zero🇫🇷🇬🇧🇵🇹🇪🇸🇩🇪[https://t.co/D7gadZ36Ib], Kyutai's latest real-time speech translation model.
I'll be giving an oral presentation on July 8 at 10:30 AM KST. Feel free to join if you'd like to learn more!💬
🎰 Welcome to the FID Lottery.
We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID.
A 2.1-point spread, pure luck. Step onto the floor 👇🧵