I needed a way to know if our model is actually better. So I built a blind test into our production platform and let 5,098 users decide.
We tested S2 Pro and S1 against ElevenLabs V3, ElevenLabs 2.5 Flash, ElevenLabs Multilingual V2, Inworld TTS 1.5 Max, and MiniMax Speech 2.8 HD.
S2 Pro ranked #1 overall, nearly 1.7x the next closest model.
But what stood out most was the per-language gap. In Chinese, S2 Pro scored 8.11. The next closest competitor scored 2.36.
Most TTS companies build for English first and then extend it to fit everything else. Our multilingual data pipeline is why the gap is widest where others invest the least.
Feel free to check out the our report at https://t.co/Pa5CFa15tY
Fish Audio performed a 10-day blind test across 5,098 real user comparisons on production traffic. Our S2 Pro model ranked #1 overall with a 66% win rate - outperforming the next closest competitor by over 60%.
- S2 Pro beat ElevenLabs V3 60% to 40%
- S2 Pro beat Inworld 80% to 20%
- S2 Pro beat MiniMax 95% to 5%
- #1 in every language tested
No one knew they were part of a test. They just picked the voice they liked more.
Check full methodology and results here → https://t.co/wFIc0oOy3x
Today we launch Fish Audio S2, a new generation of expressive TTS with absurdly controllable emotion.
- open-source
- sub 150ms latency
- multi-speaker in one pass
Real freedom of speech starts now 👇
🎉 Congrats to @FishAudio on launching Fish Audio S2, a frontier TTS model with fine-grained prosody & emotion control via natural-language inline tags. SGLang Day-0 support is now live!
🏆 Best WER on Seed-TTS Eval; 81.88% win rate on EmergentTTS-Eval
🎙️ Voice cloning with 86.4% prefix-cache hit rate via RadixAttention
⚡️ RTF 0.34, 63.3 tok/s on single H200 (single batch)
🌍 Trained on 10M+ hours of audio across ~100 languages, GRPO-aligned
🔧 Dual-AR (Slow + Fast AR) is LLM-isomorphic: continuous batching, paged KV cache & CUDA graphs inherited natively
🗣️ Native multi-speaker: turn-taking, interruptions & cross-speaker emotion in a single pass
👉Cookbook: https://t.co/Adc5aRtf4e
👉Blog: https://t.co/95qdBoiaPe
🎬 Curious how to run with SGLang? Check out this voice cloning demo from @GenAI_is_real with Fishaudio-S2-Pro:
Today we’re excited to publicly launch Fish Audio S1, the most expressive and natural TTS model on the market.
- 6x cheaper than elevenlabs
- 20K active developers
- 5M ARR (real ARR)
Clone your own voice for free and get started 👇
Introducing OpenAudio S1! 🎉
Command your AI voice actor like never before:
- 🔥 Experience unparalleled expressiveness & naturalness: Voted #1 in TTS-Arena!
- 🎯 Achieve state-of-the-art accuracy: 0.008 WER & 0.004 CER in Seed TTS Eval.
- 💬 Control a spectrum of emotions via natural language: From (angry), (happy), (sad) to nuanced (emphasize), (whispering), (empathetic), and more!
- 🎬 Unleash your creative vision: Ideal for Video, Audiobooks, Podcasts, AI Companions, Gaming & endless possibilities!
🔗 Explore OpenAudio S1: https://t.co/xgrGPGQcE8
📖 Get the full story: https://t.co/8z1TxaSlM6
We’re excited to release ACE-Step / ACE-Step-v1-3.5B, a fast, versatile DiT-based foundation model for music generation that runs on consumer-grade GPUs.
With its simple architecture and low hardware requirements, it’s easy to fine-tune for various music tasks, empowering, not replacing, artists and creators.
Think of it as a step toward music’s Stable Diffusion moment.
※ Trained on authorized, purchased data.
Demo Page: https://t.co/tVc8xCZfdO
Hugging Face: https://t.co/9sjZ5Dntvv
Git repo: https://t.co/AzHEwgmZgr
OpenAI just released AGI (all ghibli images)
Today we launch W25.
Our most powerful batch yet.
They broke the benchmarks for acceleration so much that 0-$3M will never hit the same…and if you want to know why we’re building a ring of servers around the sun, watch carefully.
If you haven’t adopted this new way of building companies, you will get outcompeted.
In twelve weeks
- four of ten teams broke $2M ARR
- fifth team hit $1.7M
- top team went from 1 to 10M
Launching applications today. Ten teams. Only one HF0.
We're deeply grateful to our team for their outstanding work in making this state-of-the-art open-source TTS possible. We're excited to see what innovations the open-source community will create with Fish Speech 1.5.
We're growing! If you're a talented sales professional or engineer interested in shaping the future of voice technology, reach out to us at [email protected].