We released Sonic-3.5 and Ink-2, the #1 streaming models for text to speech and speech to text you can use in your voice agents today.
New architectures enable new frontiers for speed and quality.
We're now the only provider to have #1 models for both speaking and listening.
New to Ink-2: keyterm prompting
Pass rare terms like product names, industry jargon, and even French words for accurate transcription. 20% higher keyword recall, with no added latency.
Learn more → https://t.co/sAphUQdlG0
Our London office is growing quickly, reach out to @_albertgu or me if you’d like to work on a very different, new and exciting research agenda to advance multimodal models
"Is this TTS model good?" gets harder to answer as models improve.
"Good" is at least five axes: correctness, naturalness, contextual correctness, robustness, and most evals only capture the first.
We wrote up the failure modes that make TTS eval hard: https://t.co/ILbJM3zXyL
Today, @cartesia and @p0 are making it possible for your voice agents to search the web at conversational speed.
Cartesia builds the fastest voice models available, and Turbo mode for Parallel Search extends that same low latency to web search. An agent built on Cartesia can query the live web without interrupting the natural flow of conversation.
Learn more 👇
For voice agents, STT has to nail three things - accuracy, turn detection, and latency. If any one falls short, the experience breaks down: the agent misunderstands, interrupts, or just feels slow.
We built Ink-2 to lead on all three. Here’s how it stacks up against other providers: link in comments.
Voice quality can be hard to evaluate objectively - preference for a given voice can confound assessment of the underlying model. @ArtificialAnlys Controlled Voice Arena addresses this by holding the voice constant (via cloning) across all models, isolating raw model quality as the variable of interest.
Under this controlled setup, Cartesia's Sonic-3.5 ranks first overall (1,122 Elo), and leads independently on both US (1,139) and UK (1,103) accents.
We're excited to keep pushing the bar on model quality for our users.
Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices
The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices.
We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice.
Key results
➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld_ai Realtime TTS-2 - Research Preview (1,070)
➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059)
➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd
➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930)
See more details below ⬇️
Cartesia Sonic 3.5 is the #1 streaming TTS model on Voice Arena US English Leaderboard.
In the overall leaderboard (streaming + non-streaming) it jumped from rank #9 → #2, moving ahead of Grok TTS, ElevenLabs v3 & OpenAI's gpt-4o-mini-tts.
Sonic-3.5 is the latest TTS model from @cartesia . It supports 42 languages, with 500+ voices available out of the box. The model has been highly preferred among raters on @voicearena_ai .
Results backed by 11,110 blind, head-to-head listener votes
Announcing the Bolna × Cartesia VOC-A-THON!
Calling the most cracked builders in Voice AI. Come ship voice agents powered by Sonic 3.5, Cartesia's most natural and expressive TTS model yet.
- Build with Sonic 3.5, Cartesia's newest TTS model
- @bolna_dev and @cartesia teams in the room
- Free Sonic 3.5 & Bolna credits for every participant
- Exciting prizes for the best voice agents
Apply now: https://t.co/n97fbUvz5N
Cartesia Ink-2 debuts as #1 for accuracy on the brand-new streaming speech-to-text leaderboard from @ArtificialAnlys! We designed Ink-2 from the ground up for voice agents - with low latency, eager transcripts, and semantic endpointing.
A great speech-to-text model for voice agents first and foremost needs to have high accuracy in production settings - this means noisy environments and conventionally difficult audio like silences, short transcripts, phone numbers, and UUIDs.
For the conversation to be smooth, it also needs to have low latency with eager transcripts to reduce end to end response time.
Finally, semantic endpointing with high accuracy is critical so they respond appropriately and don't interrupt the user.
🚨 AVTR-1 New Model is OPEN WEIGHTS . Duplex Native , #1 on benchmarks.
Here’s what being released. Links in comments
- Model + Paper now on HF
- Full Github repo to run it really fast
Run it anywhere as low as $0.
Comment, share, star on GH to get the word out
Sonic 3.5 is now the #1 text to speech model on the @ArtificialAnlys leaderboard!
You no longer have to trade off quality and latency - Sonic 3.5 also has the fastest time to first audio at 82ms end to end.
See full benchmark results 👇
Cartesia’s Sonic-3.5 takes the #1 spot on the Artificial Analysis Speech Arena Leaderboard, surpassing Inworld Realtime TTS 1.5 Max and Google’s Gemini 3.1 Flash TTS
Sonic-3.5 is the latest TTS model from @cartesia . It supports 42 languages, including 9 Indian languages, with 500+ voices available out of the box. The model has been highly preferred among voters in the TTS Arena, with its demonstrated naturalness and accurate transcript following.
Key takeaways:
➤ Quality: Sonic-3.5 has an Elo score of 1,218 (+16/-16) based on 1,144 arena appearances, placing it ahead of Inworld Realtime TTS 1.5 Max at 1,194 and Gemini 3.1 Flash TTS at 1,209
➤ Pricing: Sonic-3.5 is priced at $39/1M characters, a premium compared to Gemini 3.1 Flash TTS at $18.3/1M characters, and Inworld Realtime TTS 1.5 Max at $35/1M characters
➤ Speed: 105.5 characters per second, compared to 205 characters per second for Inworld Realtime TTS 1.5 Max and 26.3 characters per second for Gemini 3.1 Flash TTS
See more details and listen to samples below 🧵