TTS outputs finally sound human. Yet, picking the right model for your language and use case is still a guessing game. VoxClash is our shot at fixing that.
TTS outputs finally sound human. Yet, picking the right model for your language and use case is still a guessing game, because the tools built to answer that question are outdated. VoxClash is our shot at fixing that.
Text-to-speech has gotten good enough that people often can't tell a synthetic voice from a real one, at least on a simple sentence read aloud.
That changes what's actually at stake.
A few years ago a slightly robotic voice was a secret tradeoff waiting to be discovered, buried inside a bigger product. Not anymore. Now, for a voice agent or a customer support bot, the voice is the whole experience. If it sounds off, the entire experience feels broken.
While it's good news on one hand, it has exposed a chink in the armour.
If you're a company trying to pick a TTS vendor, you want one simple answer.
Which model actually sounds best for your language, your use case, your deployment.
Right now, nobody can give you that answer with any real confidence. The industry's two default answers kinda don't hold up anymore.
Let’s start with Mean Opinion Score (MOS) where listeners rate a clip from one to five on how natural it sounds. It used to separate models cleanly. It doesn't anymore.
Modern TTS systems cluster so close to human reference audio that a five-point scale can't tell them apart, even when listeners genuinely prefer one over another. The score isn't stable on its own terms either.
Swap the anchor clips or the rater pool in a session and the same model's placement shifts, so two labs running MOS on the same set of models can land on different rankings without either team doing anything wrong.Â
Most tests never specify a use case, so a naturalness rating collected for narration tells you nothing about a support call.Â
And when a model does lose, MOS hands you a number and nothing else, no sense of whether it mispronounced a name, dropped a pause, or just sounded robotic.
Word Error Rate (WER) doesn't fix this either, because it's answering a different question entirely.
WER synthesizes a clip, runs it back through a speech recognizer, and compares the transcript to the original text, so what it actually measures is partly the model and partly whichever ASR system did the transcribing.Â
Once a TTS model clears basic intelligibility, and most frontier models do, the remaining differences in WER shrink into noise well under 2%, right as listeners keep hearing real gaps in prosody and delivery. A clip can score a lower WER than the human reference recording and still sound worse to human listeners, because WER treats every valid delivery of a sentence, excited, flat, sarcastic, as equally correct once the words match.Â
None of that carries across languages either, since word counts and tokenization aren't comparable from English to Hindi to Telugu.
So the two metrics that built the last decade of TTS benchmarking are running out of runway. Meanwhile the alternatives buyers actually reach for don't agree with each other. Every vendor cites a different third-party leaderboard, run on different prompts, different voices, different languages, under protocols nobody outside that lab can fully see.Â
That's the gap.
What's missing is a leaderboard that tags what actually went wrong when a model lost. That's what VoxClash adds. A mispronounced name, a dropped pause, a delivery that sounds synthetic, the list goes on. VoxClash is built to reveal these intricacies. Â
The Case for a Different Kind of Leaderboard
The field has been moving toward the same answer for a while now.
Chatbot Arena established a template for this: let people vote between two outputs, then turn those votes into a ranking instead of relying on one absolute score. Voice Arena applied the same idea to speech.
Bradley-Terry math turns win and loss records into Elo scores, and running the results through repeated resampling produces a confidence interval alongside every rank, so you know where a model lands and how sure anyone should be about it.
VoxClash by Deccan AI builds on that same foundation, purpose-built for TTS.Â
Every rank comes from blind model-versus-model votes on the same script, never a solo naturalness rating.
Every score ships with a 95% confidence interval and a rank range.
Six issue tags attach to every clip after the vote, so a loss comes with a reason instead of just a position.
Speed runs on its own track too, separate from preference, because how fast a model responds and how good it sounds are two different questions with two different answers.
‍

‍
How VoxClash Works
A vote is simple. Two clips play, generated from the same script by two different models, and a rater picks the one that sounds better, calls it a tie, or flags both as bad.
Only decisive votes, a clear A or B, feed into the ranking math. Every model speaks through one fixed default voice throughout, so a comparison measures the model and nothing else.
Every clip carries a version tag tied to its run, so a script recorded in November reads as a distinct dataset from the same script recorded later.
Models with fewer head-to-head matchups get paired more often, so no leaderboard position rests on a thin sample.
The win and loss records feed a Bradley-Terry model, which turns them into Elo scores, and running that fit across 200 resampled versions of the same data produces the confidence interval and rank range you see next to every score.
After a vote, either clip can carry one of six issue tags, covering a mispronunciation, unnatural prosody, a robotic tone, a code-switch error, a mismatched emotion, or a cut-off or glitch.
‍
.png)
‍

‍
What VoxClash v1.0 Covers
VoxClash launches with three languages, English (India), Hindi, and Telugu, each tested across 50 scripts, 9 models total across all three. That's the starting set, not a permanent boundary. Global language coverage is next. The v1 scripts already carry code-switching passages, tagged directly in the voting interface, so a model's handling of mixed-language speech counts as part of the baseline evaluation from day one.
What's Next for VoxClash
VoxClash runs on internal hosting today. Public v1 is launching soon. From there, the roadmap adds domain-specific script packs across appointments, billing, customer support, and news anchoring, broader language coverage beyond the three launch markets, and voice-versus-human comparisons alongside the existing model-versus-model format.
VoxClash runs on internal hosting today. Public v1 is launching soon.We’ll be adding domain scripts for appointments, billing, customer support, and news anchoring. Parallely expanding past the three launch languages and adding voice-versus-human comparisons alongside the existing model-versus-model format.
.png)
%20(1).png)