Introducing Octave 2: our next-generation multilingual text-to-speech model
What’s new:
- Fluent in 11+ languages
- 40% faster (<200ms latency) & 50% cheaper than Octave 1
- Multi-speaker conversation
- More reliable pronunciation
- New voice conversion & phoneme editing capabilities
For the month of October, we’re offering 50% off our Creator plan - use code OCTAVE2 at checkout!
Excited to share my co-first paper at @Nature! I would like to thank my co-first authors Jason and Heydar and Prof. Dan Dombeck for his amazing mentorship. I hope that this work can open up future investigations of the underlying neural mechanisms of representation drift!
2024: Voice Cloning
2025: What about personality cloning?
Hume’s voice AI can now not only mimic your voice but also speaking style and language.
It’s now available via our TTS and new speech-to-speech model, EVI 3, which is also launching today.
We released our new conversational speech-llm. Fully multimodal (inputting/outputting text&audio tokens). Ability to adopt any voice, modulate speech, and generate language.
Meet EVI 3, another step toward general voice intelligence.
EVI 3 is a speech-language model that can understand and generate any human voice, not just a handful of speakers. With this broader voice intelligence comes greater expressiveness and a deeper understanding of tune, rhythm, timbre, and speaking style.
Today, we’re releasing Octave: the first LLM built for text-to-speech.
🎨Design any voice with a prompt
🎬 Give acting instructions to control emotion and delivery (sarcasm, whispering, etc.)
🛠️Produce long-form content on our Creator Studio
Unlike traditional TTS that just “reads” words aloud, Octave understands how meaning affects delivery to generate emotional, human-like speech.
Excited to share what I've been working on! Over the past few months, we’ve been aligning language and speech both semantically and non-semantically in a fully trained multimodal model, while maintaining language capabilities.
Introducing OCTAVE, a next-generation speech-language model.
OCTAVE has new emergent capabilities, like on-the-fly voice and personality creation and much more 👇
Hume AI, a startup founded by a psychologist who specializes in measuring emotion, gives some top large language models a realistic human voice. https://t.co/BASqQRBbGE
Introducing Empathic Voice Interface 2 (EVI 2), our new voice-to-voice foundation model. EVI 2 merges language and voice into a single model trained specifically for emotional intelligence.
You can try it and start building today.
We’ll probably look back on Her as the most accurately predictive sci-fi movie of our generation.
Personal (even intimate), long-context, always-on AI assistants are happening.
It feels like https://t.co/uj9nfZSStQ this week filled the missing piece of the puzzle.