Gemini (Duplex) is the most human-sounding AI voice agent in #Indian English - 1,101 Elo across 1,543 blind human votes.
It still trails a real human by 223 Elo.
Elo is a rating system used to rank performance in head-to-head matchups. Here, it measures how often an AI voice is judged more human than another voice.
Voice of India's Speech-to-Speech Arena scores full-duplex voice agents on humanness: two agents hold a natural conversation, blind listeners pick which one feels human. One of the five entrants is an actual person: the Human Anchor.
Here’s where the board stands:
➤ Human Anchor: 1,324 Elo, wins 90% of its matchups. The nearest AI sits 223 Elo back.
➤ Gemini (Duplex)- @GoogleDeepMind: 1,101 Elo, 65% win rate. The top machine, +204 clear of the next.
➤ OpenAI (Duplex)- @OpenAI: 897 Elo, 36% win rate.
➤ Grok (Duplex)- @xAI: 844, and Inworld (Duplex)- @inworld_ai: 833. Inside each other's error bars (±22–23); ranks 4 and 5 interchange.
The gap that matters here isn't between the models. It's the 223 Elo between the best AI and one human voice. That's a line no other voice arena can draw, because no other arena puts a human on the board.
Bradley-Terry Elo from 1,543 blind pairwise votes, 14% ties. Provisional, and moving as votes come in.
Full board, and cast your own vote → https://t.co/MAUHOLC0dx
Congrats to @GoogleDeepMind on the top machine slot, and to @OpenAI, @xAI and @inworld_ai for stepping into a human-anchored arena.
Nano Banana 2 is the most culturally accurate image model on Indian prompts.
1,112 Elo across 22,865 blind human votes.
Elo is a rating system used to rank performance in head-to-head matchups. Here, it measures which model produces the more culturally accurate image when people judge the outputs blind.
And even the leader loses roughly a third of its matchups to models ranked below it.
Voice of India's Text-to-Image Leaderboard scores cultural truth, not aesthetics: 1,000 prompts written by people living in each state about something specific to where they live, then voted blind by raters from that same state. The people who notice if the drape, the sweet, the script or the temple is wrong.
➤ Nano Banana 2- @GoogleDeepMind: 1,112 Elo, 67% win rate. The leader, +45 clear of the next.
➤ Grok Imagine- @xai: 1,067 Elo, 61% win rate.
➤ GPT Image 2 High- @OpenAI: 1,039 Elo, 56% win rate.
➤ FLUX.2 Max- @bfl_ml: 993, MAI Image 2.5- @MSFTAI: 985, and Krea 2 Medium- @krea_ai: 983. Inside each other's error bars (±6 to 7); ranks 4 to 6 interchange.
➤ Recraft v4.1 Utility- @recraftai: 914, 36% win rate, and Reve 1.0- @reve_image: 907, 35%.
The gap that matters here isn't between the models. It's between a beautiful image and a real one.
Every system on this board can render a gorgeous Indian street; 205 Elo separates them on whether it's the street the prompt actually asked for, or the default sepia Rajasthan lane a web corpus taught them to draw.
Bradley-Terry Elo from 22,865 blind pairwise votes, ties counted as half a win each, intervals from resampling the prompt set. Provisional, and moving as votes come in.
Full board, state-wise breakdowns, and cast your own vote → https://t.co/MAUHOLBsnZ
Congrats to @GoogleDeepMind on the top slot, and to @xai, @OpenAI, @bfl_ml, @MSFTAI, @krea_ai, @recraftai and @reve_image for putting their models on a board that scores accuracy over polish.
And this is just Hindi, the most widely spoken language in India, the ‘easy’ case, where every major model still scores under 12%.
The full benchmark: 15 languages · 139 regional clusters · 536 hrs · 36,691 speakers · 306,230 unscripted utterances. Closed- so no one can train on the test set.
Full board + methodology → https://t.co/uUyOc05fjv
Built with @AI4Bharat (IIT Madras)
Which speech-to-text model understands Hindi best?
Take a look.
We tested every major system on 21,020 real, unscripted #Hindi utterances, scored with OI-WER - the metric that recognizes different variations of Hindi, including code-mixed Hindi, rather than unfairly penalising them as errors.
#1 is Made in India: @sarvamai's Saaras v3 (3.78%) - ahead of @Google, @Amazon, @Microsoft, @Meta and @OpenAI.
Live board 👇 https://t.co/MAUHOLC0dx 🧵
Why OI-WER?
Indians code-switch constantly, and one code-mixed word has many valid spellings. Standard WER counts those as errors- making every model look worse than it really is.
OI-WER (Orthographically-Informed WER) accepts every human-verified spelling. Models are judged for what they heard, not how they spelled it.
Announcing the https://t.co/MAUHOLC0dx TTS Leaderboard, benchmarking how text-to-speech models perform on naturalness across Indian languages when judged by native listeners. We are initiating coverage with Maya Research, Google DeepMind, Cartesia, Gnani, ElevenLabs, Smallest AI, Sarvam AI, Google, MiniMax, OpenAI, and Microsoft.
TTS is the output layer of every voice agent, IVR, and accessibility product shipping in India. Vendors publish naturalness claims, but those claims are almost always measured in English, or on read studio speech. Teams deploying in Hindi, Tamil or Odia have had nothing to check them against. We are extending human evaluation to Indian languages, so buyers can pick a TTS provider on measured naturalness in the language they actually ship in.
Each result pairs a model with a fixed voice, selected once per model per language and held constant across every battle, so a vendor's default voice choice does not decide the ranking. Only the model behind the audio changes.
At launch, the leaderboard covers 14 models across 10 languages plus a Finance (Hindi) domain board, and we will keep expanding coverage as more systems add Indian language support.
Key elements of the https://t.co/MAUHOLC0dx TTS Leaderboard:
➤ Blind pairwise comparison, not MOS: modern TTS systems score statistically indistinguishable from human speech on five-point scales, so absolute ratings no longer separate frontier models. Raters hear the same sentence from two systems back to back and choose, with Both and Neither available so ties and shared failures are captured rather than forced into a winner
➤ Screened native raters: candidates pass an auditory discrimination test, then justify their choices against the perceptual criteria before receiving training and access to real work. Each rater evaluates 150 randomly sampled sentences, so no single ear moves a board
➤ Natively authored sentences across 16 deployment domains, in three input conditions: normalized (numerals expanded), symbolic (raw numerals and formulas as they arrive from upstream systems), and code-mixed (English inside Indian-language sentences)
➤ Bradley-Terry Elo with bootstrapped 95% confidence intervals: models whose intervals overlap share a rank and are reported as tied. Every language board is fit independently and no cross-language aggregate is published, because scripts and morphology differ
Key results (Hindi, General):
➤ The top three are a statistical tie: Maya 2 Native (1,079 ±17) @mayaresearch_ai , Gemini 2.5 Pro TTS (1,076 ±9) @GoogleDeepMind , and Sonic-3.5 (1,072 ±21) @cartesia all carry rank ranges spanning position 1. We are not calling a winner on Hindi
➤ Four Indian systems placed in the top 9: Maya 2 Native at 1, Vachana TTS v3 at 5, Lightning 3.1 Pro at 7 and Bulbul v3 at 9. Frontier multilingual models and Indian-language specialists run on the identical corpus with no separate track
➤ Frontier US models cluster at the bottom on Hindi: GPT-4o Mini TTS ranks 12th (972), MAI Voice 1 13th (943) and Azure TTS 14th (921)
➤ Most of the board is not separated: the spread from 1st to 8th is 70 Elo points across 14 models, narrower than several of the confidence intervals. Positions 4 through 11 are largely interchangeable and should not be read as a ranking
Full Hindi board, 14 models, 21,770 blind votes from native listeners.
Rank ranges are published alongside rank. Maya 2 Native holds rank 1 but its range spans 1 to 3, and Sonic-3.5 at rank 3 spans 1 to 4. A model's position is only meaningful to the width of its interval.
Win rate is reported separately from Elo. Azure TTS wins 38% of its head-to-head matchups in Hindi. The board leader wins 59%.
Introducing Voice of India, India's first independent, multimodal AI evaluation platform, built in partnership with AI4Bhārat at IIT Madras.
Shri S Krishnan, Secretary, MeitY, Government of India, joined us in Delhi for the launch.
One line from the conversation stayed with us. “You cannot build AI for India without building the ability to measure it for India.”
India has 22 official languages, hundreds of dialects and highly diverse ways of communicating. Yet most AI systems don't reflect the the reality of how people actually speak and communicate in India.
That is the gap we are building for. Results will be published here.
Stress tested Sarvam AI, India's sovereign TTS model through a Large-Scale Indic TTS Evaluation of @SarvamAI's Bulbul V3 vs @elevenlabs's v3 Alpha and 2.5 Flash vs @cartesia's Sonic-3 with over 44,000+ votes from 1000+ humans.
More on this in the next few days with audio samples where each model excelled vs the others.
Will AI take away our jobs or transform the way we work?
In Ep 1 of The Skill Edit, our guests reflect on a question weighing heavily on the minds of today’s youth:
Is technology here to replace us…or elevate us?
🎧 Full episode out now. Link in the comments!
#workforce#AI