New model OuteTTS-1.0-0.6B
- Based on Qwen-3 0.6B
- Apache 2.0 (free for commercial and personal use)
- Supports 14+ languages
- Adds batched inference, enabling fast audio generation for long inputs (~0.1โ0.02 RTF)
๐ฆ Model Weights: https://t.co/gn9dqNDNQc
@JUQ_AI Yes, you can definitely add a new language by further pre-training the model or fine-tuning it. Check out the Unsloth guide: https://t.co/O5kAgswekl "Oute-TTS (1B)"
OuteTTS v1.0: A multilingual speech synthesis + voice cloning model built on Llama3.2-1B.
Supports one-shot cloning (~10s audio), native text input (no romanization), and automatic word alignment.
Compact, fast, and fully local - runs on llama.cpp or EXL2. @OuteAI
OuteTTS 1.0: Upgrades in Quality, Cloning, and 20 Languages
- Best Performance: Generate audio around 42 seconds in a single run (approximately 8,192 tokens). It is recomended not to near the limits of this windows when generating. Usually, the best results are up to 7,000 tokens.
- Context Reduction with Speaker Reference: If the speaker reference is 10 seconds long, the effective context is reduced to approximately 32 seconds.
Github-link in comments
New OuteTTS model is here Llama-OuteTTS-1.0-1B
Major upgrades in speech synthesis & voice cloning, now with smoother processing and native multilingual support for 20 languages!
๐ Full details on what's new & weights:
https://t.co/akF1QwgG8w
Introducing OuteTTS 0.3 1B & 500M ๐ฅ
> Zero shot voice cloning
> Multilingual (en, jp, ko, zh, fr, de)
> Trained on 20,000 hours of audio
> Powered by OLMo-1B & Qwen 2.5 0.5B
> Speed & emotion control
> Powered by HF grants ๐ค
Open science ftw!
Introducing TTS WebGPU: The first ever text-to-speech web app built with WebGPU acceleration! ๐ฅ
High-quality and natural speech generation that runs 100% locally in your browser, powered by OuteTTS and Transformers.js.๐ค Try it out yourself!
Demo + source code below ๐
Smol TTS keeps getting better! Introducing OuteTTS v0.2 - 500M parameters, multilingual with voice cloning! ๐ฅ
> Multilingual - English, Chinese, Korean & Japanese
> Cross platform inference w/ llama.cpp
> Zero-shot voice cloning
> Trained on 5 Billion audio tokens
> Qwen 2.5 0.5B LLM backbone
> Trained via HF GPU grants
Model weights on the hub, you can even run this on a Raspberry Pi! Go run, inference now! ๐
@harambe_musk Not at the moment, but thereโs nothing too special about the training. Itโs basic pre-training, similar to how you would train any other language model for next token prediction.
๐ถ OuteTTS-0.1-350M: Teaching Language Models to Speak Using Audio Tokens and Forced Alignment!
Our new experimental text-to-speech model with only 350M parameters achieves high-quality speech using a pure language model-based approach.
https://t.co/CmgOtVVroS