This is what happens when you plug LLMs into voice assistants, instead of a decade of handwritten rules.
This video dissects Voxtral (a family of OSS speech models) and the foundational work behind it (audio tokenization, semantic/acoustic disaggregation, etc).
Thank you @MistralAI for your collaboration and for your detailed technical reports in an increasingly opaque industry!
00:00 Intro
01:03 Modular vs end-to-end speech models
03:30 Speech-to-Text
06:07 Delayed Streams Modeling (DSM)
09:41 Whisper Streaming
10:33 Voxtral Realtime
13:07 Voxtral Text-to-Speech
14:28 Throwback: WaveNet
15:24 Audio tokenization
20:39 The Voxtral Codec
21:49 Back to Voxtral TTS
25:30 Outro
Check out our Voxtral paper now on arxiv: https://t.co/PCOJ3e9XJn
Details on on pre-training, fine-tuning and alignment, with ablations covering how to chose the optimal model architecture and pre-training format!
In our continued commitment to open-science, we are releasing the Voxtral Technical Report: https://t.co/CqZSc2sMu3
The report covers details on pre-training, post-training, alignment and evaluations. We also present analysis on selecting the optimal model architecture, which pre-training format to use, and the benefits of DPO.
Check out blog post - https://t.co/htZXHafKbS
HuggingFace:
- Voxtral Small - https://t.co/1vJEs4Zrkb
- Voxtral Mini - https://t.co/ttyC27vxt7
Lot more come soon ...
I've been working on speech recently and excited to announce Voxtral by @MistralAI - our first set of open (Apache 2.0) audio models focused on transcription and speech understanding 🎤
We're releasing two open frontier models in this category - Voxtral Mini and Small
Performance:
- SOTA transcription on variety of languages. Frontier performance-to-price
- Outperforms models in its weight category on understanding and chat metrics
- Retains strong text performance and serves as a drop-in replacement for Mistral Small 24B and Ministral 3B
Introducing Mistral Small 3.2, a small update to Mistral Small 3.1 to improve:
- Instruction following: Small 3.2 is better at following precise instructions
- Repetition errors: Small 3.2 produces less infinite generations or repetitive answers
- Function calling: Small 3.2's function calling template is more robust
📰 News in Arena: Mistral Medium 3 makes a strong debut with the community!
Highlights:
💠 #11 overall in chat: a +90 point leap from Mistral Large
💠Top-tier in technical domains (#5 in Math, #7 in Hard Prompts & Coding)
💠#9 in WebDev Arena
Congrats to @MistralAI on the solid performance. 👏 Can’t wait for the next Mistral Large!
Bard will undoubtedly eat ChatGPT's lunch. The product is a lot nicer to use already and their teams are iterating super fast!
According to some people at OpenAI, the org is in the worst of both worlds, too ossified to move like a start-up, and too small for big co moves.