Announcing the Voice Arena STT Leaderboard, our human-verified benchmark testing how accurately speech-to-text models transcribe real, spontaneous conversation across 7 languages
The leaderboard covers ๐บ๐ธ US English, ๐ฎ๐ณ Hindi, ๐ง๐ฉ Bangla, ๐ง๐ท Brazilian Portuguese, ๐ป๐ณ Vietnamese, ๐ท๐ด Romanian and ๐ธ๐ฆ Arabic. Every model is evaluated on the same held-out corpus per language: real, unscripted phone calls between native speakers, recorded in-country on their own phones, with the noise, fillers and interruptions of natural conversation.
Unlike public STT benchmarks, results are measured on fully proprietary audio - no model has seen a second of it in training. @Reson8Offical's resonant-1 leads the US English board at 4.34% WER, ahead of models from @Microsoft , @Google , @OpenAI and @Meta . @SarvamAI leads in Hindi and Bangla, while Microsoft takes #1 in Romanian and Brazilian Portuguese.
Key elements of the Voice Arena STT Leaderboard:
โค Real speech, real conditions: audio is drawn from unscripted phone calls between native speakers, recorded in-country on the speakers' own devices, preserving the background noise, fillers and interruptions of production traffic.
โค 100% proprietary, held-out audio: the corpus is collected and owned by Voice Arena. No model has seen a second of it in training, so results cannot be inflated by dataset contamination.
โค Human-verified ground truth: transcripts are produced by native transcribers through a six-stage pipeline, and only segments verified 3-of-3 survive into the final test set.
โค Corpus-level WER with symmetric, language-aware normalization: the same normalization is applied to both references and hypotheses, so models are penalized for recognition errors, never for spelling choices.
Key results for the Voice Arena STT Leaderboard:
โค The most accurate US English model is not from a big lab: Reson8's resonant-1 leads at 4.34% WER, ahead of Microsoft, Google, OpenAI and Meta.
โค Model choice alone is worth ~44%: on the US English board, #1 sits at 4.34% WER and #20 at 7.74%. Same audio, same test set. Picking the right model cuts nearly half of your transcription errors before you touch anything else in the pipeline.
โค There is no global winner: the #1 model changes with the language, with Reson8 leading US English, Sarvam leading Hindi and Bangla, and Microsoft leading Romanian and Brazilian Portuguese. Teams building for more than one market will not be covered by a single vendor at the top of every board.
Your vocabulary. Your domain. Your language.
We raised โฌ5M to build speech recognition that customizes in seconds.
Investment by @balderton , with participation from @nphardvc.
We're hiring โ https://t.co/Y0V8dDUWVk
Try us โ https://t.co/D1moI2VUSO
Amsterdam-based @Reson8Offical, a speech AI startup building hyper-customised automatic speech recognition (ASR), has raised a โฌ5 million pre-Seed round to scale its Europe-based infrastructure. ๐ณ๐ฑ ๐๏ธ๐ค
@balderton
https://t.co/lCXLozT553