Today we're releasing Soniox v4 Real-Time, our most advanced real-time speech recognition model yet, and a new standard for real-time speech AI.
Soniox v4 Real-Time delivers speaker-native accuracy across 60+ languages with industry-leading low latency, purpose-built for voice agents, live captioning, and mission-critical voice applications.
What’s new in v4 Real-Time 👇
🌍 Speaker-native accuracy for 60+ languages
We've ended the "English-first" era. Major and minor languages now get the same level of accuracy in a single unified model.
⚡ Millisecond finality
High-accuracy final transcripts arrive just milliseconds after speech ends, enabling faster LLM responses and more natural conversations.
🧠 Semantic endpointing
The model understands intent, not just silence, resulting in fewer interruptions, smoother turn-taking, more human-like voice interactions.
🌐 Real-time global translation
Transcribe and translate simultaneously in one real-time stream, with low-latency streaming translation in chunks (not sentence-level) as speech happens.
Upgrading is seamless
Soniox v4 Real-Time is fully backward-compatible with v3 Real-Time. Switch to stt-rt-v4 and immediately benefit from higher accuracy, lower latency, and improved reliability.
👉 Read the full announcement https://t.co/dxIJwz3uzD