@kirtandopamine AI can do infra, check naive. The problem is when AI is spinning more than a lightweight backend frontend server then it automatically becomes expensive which goes beyond hobbyists and vibe coders. And experienced people already use claude to spin up the infra.
@sridharfyi We are building https://t.co/hGVFAshzS7 Please don’t mind the lousy content copy. Try out the app, the tech works well. We need GCP credits if possible or any other platform would also work.
Wow!
I am testing this FREE open source text to speech AI system and I am blown away.
Can clone any voice in less than a few minutes of speech.
@resembleai has changed the entire game.
Thank you!
I’ll write more about it soon.
“If this is the best model, why open source it?”
Because the goal isn’t to fight closed-source companies. We want them to succeed. This is a massive market. We can all win.
The real advantage of open source isn’t ideology. It’s adoption.
Not everyone can (or should) pay $50k/month to run voice models. A team in Southeast Asia shouldn’t have to choose between latency, cost, and quality.
With open models:
- You can run local or regional inference
- Pay a fraction of the cost
- Get lower latency
- And keep the same output quality
That changes who gets to build.
And it’s not just builders, it’s researchers too.
Before today, working with voice models meant hitting a wall. No access, no internals, no real experimentation.
Now?
You can break the model.
Rebuild it.
Run your own evals.
Optimize inference.
Push the field forward.
That’s how research compounds.
That’s how ecosystems grow.
Open source isn’t about winning a war.
It’s about unlocking the next wave of voice AI.
New TTS banger: Chatterbox Turbo 🤯
Zero-shot model that matches any reference voice with native paralinguistic tags, optimized for low-latency voice agents.
⬇️ Demo available on Hugging Face
ElevenLabs has officially LOST to Open-Source
ResembleAI allows you to clone ANY voice without verification using on 5-10 seconds of audio, and dominates on paralinguistic tags for human-like expressions.
Most "fast" text-to-speech models sound robotic. Most "quality" TTS models are slow. None incorporate authentication at a foundational level. @resembleai solved all three.
Chatterbox Turbo delivers:
🟢<150ms time-to-first-sound
🟢State-of-the-art quality that beats larger proprietary models
🟢Natural, programmable expressions
🟢Zero-shot voice cloning with just 5 seconds of audio
🟢PerTh watermarking for authenticated and verifiable audio
🟢Open source – full transparency, no black boxes
Try it on HuggingFace: https://t.co/cPXPQyPrRN
This is the DeepSeek moment for Voice AI.
Today we’re releasing Chatterbox Turbo — our state-of-the-art MIT licensed voice model that beats ElevenLabs Turbo and Cartesia Sonic 3!
We’re finally removing the trade-offs that have held voice AI back.
Fast models sound robotic. Great models are slow. And none are built for trust. We fixed all three. Chatterbox Turbo is transparent, auditable, and built for a world that needs proof.
Chatterbox Turbo is available on Replicate
• Speeds up to 6x faster than real time
• Expressive sound tags like sighs, laughs, and coughs
• PerTh watermarking on every output
https://t.co/FtOsG7FJ2i
Voice AI just crossed an important line.
Chatterbox Turbo is open-source…
and it outperforms ElevenLabs, Cartesia, and VibeVoice in blind tests.
Faster than most “real-time” models.
More expressive than most “high-quality” ones.
Here’s what’s actually new 👇
This is Chatterbox Turbo by @resembleai
Independent blind evaluations show:
• 65% win rate vs ElevenLabs
• Wins against Cartesia Sonic 3
• Wins against VibeVoice 7B
Not a demo advantage.
Not cherry-picked samples.
We raised $13M and launched DETECT-3B Omni today.
First multimodal deepfake detector ready for production.
#1 on Speech DeepFake Arena and DFBench.
98% accuracy, 40+ languages.
Backed by @Google's AI Futures Fund, @okta, @Sony_Innov_Fund, Taiwania, @kddipr, Gentree, @iagcapital, @BFFfrontierfund and existing investors at @JavelinVP, @craft_ventures, @UbiquityVC , @ComcastVentures.
Google AI Future Fund, Okta Ventures, Sony Ventures, Taiwania, KDDI.
Deepfakes caused $1.56B in fraud this year. Our investors and customers know the threat is here threat is here.
Every organization will need AI-native defense. We’re securing generative AI from creation through distribution.
Resemble AI just dropped Chatterbox, their first open-source TTS model that:
> Outperforms ElevenLabs in side-by-side tests
> Supports emotion exaggeration control for expressive speech
> Delivers ultra-low latency (<200ms) for real-time applications
> Includes imperceptible neural watermarking for responsible AI
> Is built on a 0.5B Llama backbone
> Trained on 0.5M hours of clean data
@TeksEdge@resembleai Hey David,
Probably it was just cached since it was not public earlier, it should work now. Here is our repo URL:
https://t.co/w44tbHptHR
🎉 After 2 years in production serving millions of requests, we're open sourcing Chatterbox - our state-of-the-art TTS model that just beat ElevenLabs in blind evaluations.
In recent testing, 63.75% of listeners preferred Chatterbox over ElevenLabs. Not only is it free and open source (MIT license), it's demonstrably superior to the proprietary models.
pip install chatterbox-tts
🚀 Try it now: https://t.co/dbnIft3Igy
🎧 Listen to Samples: https://t.co/VWLWBzjyyy
🤗 Hugging Face: https://t.co/xSRWIB8nyG
📊 Podonos: https://t.co/XOunvW4cVm
🚀 Big news! Our audio ‘Edit’ feature is now live! Say goodbye to those awkward [BEEP]s and endless re-recordings. Just highlight, type, and watch your audio transform 💼💪 Try Edit Now: https://t.co/YspLOi2X2z