Vapi has selected Soniox as the speech-to-text provider in three of its four new voice agent presets: Balanced, High Intelligence, and Cost Saver.
@Vapi_AI processes more than one million calls per day and has deep visibility into how speech models perform across real production traffic.
Its new Model Intelligence evaluates providers using metrics from real Vapi deployments. Soniox is listed with:
• 300 ms latency
• 1.8% word error rate
• $0.004 per minute
This is what voice AI evaluation should look like: real workloads, real production data, and metrics that matter when building at scale.
Do not trust small, closed, artificial benchmarks. Test models on real speech and real applications.
We are proud to be the speech-to-text foundation behind voice agents built on Vapi.
No keyboard nearby? No problem.
Phone + KeyMod = a pocket-sized local control console for quick terminal commands and maintenance tasks.
No Wi-Fi. No remote desktop. Just plug in and control.
#KeyMod#ServerRoom#Homelab#Sysadmin#USBHID
Don’t trust STT benchmarks.
Too clean.
Too closed.
Too English-heavy.
Too easy to cherry-pick.
Real speech is chaos: accents, noise, code-switching, interruptions, names, numbers, IDs, domain terms.
So we built Soniox Compare STT.
Open source. Raw output.
Trust no one. Test everyone.
https://t.co/ft7MpbwEki
@soniox_ai It's even more powerful when you combine realtime multilingual transcription with translation. For example an interview with Luka Dončić, where he speaks his mother tangue and English and both get translated to Italian on the fly. Works with async too.
Fair point, but for a really solid setup you will want to rely on both, VAD for interrupt detection since it fires super-fast and stt for validation, if a stop is actually needed, otherwise you will stop at "mhm" and other filler words. Either way, feedback noted we will look into it.