JUST IN: Meta just introduced Voicebox!
This is the first generative AI model that can synthesize speech across six languages, perform noise removal, edit content, transfer audio style & more.
Highlights
▸ Generalizes speech generation across tasks with impressive results and efficiency.
▸ Generates high-quality clips from scratch or modifies samples in various ways.
▸ Enables modifications to any part of speech, not just the end.
▸ Faster and surpasses existing models in speech generation quality and speed.
▸ Achieves superior word error rate and audio similarity compared to other models.