Today we’re introducing MIRA, a new multiplayer world model, built with @gen_intuition, in collaboration with Epic Games.
We release an in-depth technical report, dataset, as well as an online demo that you can try right now (link below).
With Invincible Voice, we help people living with ALS communicate more easily.
Encountering Olivier Goy, an entrepreneur who lives with ALS and relentlessly fights to help all patients, made it obvious that our cutting-edge voice AI should help.
We turned our Unmute voice-wrapper into a new system that 1/ transcribes interlocutor’s speech in real time, 2/ suggests various relevant responses via a personalised language model, 3/ utters patient's chosen response with their voice (using 10s pre-disease speech recordings).
True to our philosophy, we open-source Invincible Voice, so that developers can refine the prototype, port it from French to other languages, adapt it to other conditions (aphasia, neurodegenerative diseases) and turn it into a deployable product. Its modularity also allows it to leverage technologies developed by @GradiumAI that supports Invincible Voice by granting it free access to its multilingual speech models.
🏠 Introducing CASA: a new way to input visual information into LLMs. The current default to do that is by inserting image tokens into the text stream, but when using many images in long conversations, this floods the context window and is thus impractical for streaming inputs.🧵
Announcing Gradium, a missing link from our research to a broader audience.
In the two years since Kyutai launched, we’ve shown how our laser-focused team was able to lead innovation for cutting-edge speech models 🎙️.
Our groundbreaking open science contributions helped propel the take-off of Gradium, a startup providing industry-grade building blocks to power the next-generation of natural voice agents.
This is an important step towards building a whole and sustainable AI ecosystem in Paris, France, and Europe 🇪🇺.
Kyutai is still busy at work 🫶 We are developing world models with @gen_intuition, and on the speech front, we’ll have something big (or should we say small?) to show you soon 🎄.
Kyutai Speech-To-Text is now open-source! It’s streaming, supports batched inference, and runs blazingly fast: perfect for interactive applications.
Check out the details here: https://t.co/bQMP56XaKC
Talk to https://t.co/1ZcGtCwvgx 🔊, the most modular voice AI around. Empower any text LLM with voice, instantly, by wrapping it with our new speech-to-text and text-to-speech. Any personality, any voice. Interruptible, smart turn-taking. We’ll open-source everything within the next few weeks.
Have you enjoyed talking to 🟢Moshi? Have you dreamt of making your own speech to speech chat experience🧑🔬🤖 ? It's now possible with the moshi-finetune codebase! Plug your own dataset and change the voice, the tone and the personality of Moshi 💚🔌💿. Here's an example after finetuning w/ only 20 hours from the public DailyTalk dataset. 🧵
Meet MoshiVis🎙️🖼️, the first open-source real-time speech model that can talk about images!
It sees, understands, and talks about images — naturally, and out loud.
Voice interaction with a compact model endowed with visual understanding opens up new applications, from audio description for the visual impaired to visual access to information.
Try it out 👉 https://t.co/adcXm2Kr3J
Blog post 👉 https://t.co/3P5Y7crHBz
If you want to work on cutting-edge research, join our non-profit AI lab in Paris 🇫🇷
Thanks to Iliad Group, CMA-CGM Group, Schmidt Sciences — and the open-source community.
Meet Hibiki, our simultaneous speech-to-speech translation model, currently supporting 🇫🇷➡️🇬🇧.
Hibiki produces spoken and text translations of the input speech in real-time, while preserving the speaker’s voice and optimally adapting its pace based on the semantic content of the source speech.
Based on objective and human evaluations, Hibiki outperforms previous systems for quality, naturalness and speaker similarity and approaches human interpreters. 🧵
🚨New Work Update🚨
"How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations" to be presented at upcoming @iclr_conf - jointly done with @MoritzBoehle, @FrancescoLocat8 and Bernt Schiele.
🌟Our findings uncover a new crucial aspect of XAI🌟
#ICLR2025
Our work on transforming existing neural networks to be inherently interpretable has been accepted at #NeurIPS2024! This is joint work with @shrebox, @MoritzBoehle, and Bernt Schiele at @cvml_mpiinf.
(1/4)
Had a great time attending #ECCV2024, co-organizing a workshop, presenting two posters, and meeting a lot of people!
Thank you to everyone who visited our posters for the fruitful discussions and feedback, and to my collaborators @SwetaMahajan1, @Amin_Prc, and @MoritzBoehle!
Meet Moshiko and Moshika, the open source Moshi models 📖🟢. Moshi is a 7B text-audio model, capable of doing full-duplex conversations: it can listen and speak at any time. Plus, its inner text monologue improves the generation 💬 All on device🧑💻
🔎https://t.co/pfPlLOxBoN
Happy and excited to have two papers accepted at #ECCV2024!
1. “Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery” with @SwetaMahajan1, @MoritzBoehle, and Bernt Schiele at @cvml_mpiinf.
(1/7)
The 1st workshop on “Explainable Computer Vision: Where are We and Where are We Going?” at #ECCV2024 is accepting submissions across topics in XAI in computer vision!
The submission deadline for the proceedings track is July 24, 2024 23:59 CEST (in ~3 days!).
@eccvconf
(1/3)
Excited to present our work on Studying How to Efficiently and Effectively Guide Models with Explanations at @ICCVConference ! This is joint work with @MoritzBoehle, @Amin_Prc, and Bernt Schiele from the Max Planck Institute for Informatics.
Paper: https://t.co/xeBhhncfWj