Needs: a machine checking every sentence, and native ears confirming what it flags.
No score tells you which voice a Tamil or Telugu speaker will trust. And a voice that wins a support call can lose a collections call.
Needs: native speakers judging blind, at scale, on the lines an agent actually says.
A dataset built from real world scenarios runs down two lanes at once. An LLM pipeline checks every sentence; a human pipeline puts native speakers on the high-risk cases. Both land in the same findings, which go back to the lab for the next round.
Accent switching mid-sentence, from a regional language into Hindi, and Indian place names read wrong. Only humans hear these.
263 confirmed errors in 690 Indian English sentences. 70% were serious enough that a customer would act on wrong information.
The global English arena ranks the lab 44th of 90, judged by US and UK listeners on English prompts. On real Hindi lines, native speakers had the lab beating ElevenLabs v3 95 to 54 after a tie on demo text, and its voices top the Kannada and Telugu ladders. Native speakers are not a nicer benchmark. They are the only one that hears what the customer hears.
Two of the sixteen Hindi voices kept after screening, shown the way they appear in the lab's dashboard, down to the sentence. Names removed. Reviewer notes are verbatim; "a" and "b" are the blind positions in the pair.
The strongest female voice in the pool, firm and even from the first word to the last. She won 83% of her BFSI pairs, and reads slightly less warm on plain conversation.
First pick for BFSI and any call where the agent has to sound in control.
MONSOON2 in the Alphanumeric code line, scored bad, tagged pauses / rhythm.
24x7 in the Written shorthand (24x7) line, scored bad, tagged pronunciation, unnatural tone.
Written shorthand (24x7) line scored bad at screening, tagged pronunciation.
Lost 1 blind pair on pronunciation.
Scripted call, turn 3: "thik hai" (pronunciation).
Empathy line: reviewers split 1 and 3 of 3, one tagged wrong emotion.
Lost 3 blind pairs on unnatural tone.
Scripted call, angry, frustrated caller: “agent doesn't seem genuinely sorry and gave weird response”
Scripted call, worried, anxious caller: “Agent is not empathetic at all”
Warm, quick and natural, the most human-sounding voice when the script loosens up: she won 89% of her conversational pairs and reviewers called her natural even when she was fast. Weaker on firm BFSI asks, where the same warmth reads as soft.
Conversational and CX. Not the voice for a collections call.
24x7 in the Written shorthand (24x7) line, scored bad, tagged pronunciation.
Written shorthand (24x7) line scored bad at screening, tagged pronunciation.
Lost 1 blind pair on pronunciation.
Inform / amount line: one reviewer scored it bad, tagged unnatural tone.
Lost 1 blind pair on unnatural tone.
Scripted call, turn 4: "wasn't apologetic enough" (wrong emotion).
Scripted call, worried, anxious caller: “some gramatic mistakes”
Scripted call, irritated, no intent to pay: “Agent is sighing too much”
Send us the lines your agent actually says. We come back with a ladder, the errors that matter, and a retest date.