//  VoXClash

Blind pairwise evaluation across languages.

Model-vs-model Elo rankings, human reference anchors, and per-language quality metrics for state-of-the-art TTS systems.
Hindi
English
Telugu
Methodology
VoxClash ranks text-to-speech systems using blind human preference. Listeners compare clips without knowing which model produced them. Rankings reflect what people prefer in head-to-head listening tests, not automated scores alone.
Evaluation modes
Model rankings (live today). Listeners hear two clips from the same script, one from each of two TTS models, and choose which sounds better. Aggregated choices produce the Elo leaderboard.
Human reference comparison. The same blind setup, but one clip is synthetic and one is a native human recording. This produces a separate vs-human score and does not change model-vs-model Elo.
Issue tags
After each vote, listeners can flag specific problems (mispronunciation, unnatural prosody, robotic tone, code-switch errors, and others) separately for each clip. Tags are recorded against the model once identities are revealed. They help explain why a model won or lost; they do not directly change Elo.
Fair generation setup
• Every comparison uses the same script text for both clips, with no extra prompts or hidden instructions.
• Each model uses a fixed default voice for that language, set through the provider API.
• Audio is generated in versioned runs so rankings can be traced to a specific corpus and model snapshot.
Elo rankings and uncertainty
Model scores come from a Bradley-Terry model fit over decisive model-vs-model votes, the same statistical family used by leading AI evaluation arenas.
95% confidence interval. We estimate uncertainty with a 200-resample bootstrap: refit Elo on random subsets of the vote history and take the 2.5th and 97.5th percentiles. The leaderboard shows this as [+a, −b], the range the score could reasonably move up or down.
Rank ranges. When bootstrap resamples place a model in several adjacent ranks (for example 2nd-3rd), we show that as 2(2-3), following the same convention as Voice Arena. New votes update Elo, confidence intervals, and rank ranges on the next leaderboard refresh. Historical votes are never altered.
Latency (TTFA)
Latency is measured separately from preference. A fast model is not automatically ranked higher on quality.

We report perceived time-to-first-audio (TTFA) from a controlled probe: time until the first audio byte arrives, plus leading silence at the start of the saved clip, approximating when a listener would first hear speech.
P50 (leaderboard column): Typical wait; median across probe scripts.
P25-P75: How consistent the wait is.
Min-max: Best and worst probe runs.
Silence padding: Hush before speech begins in the clip.
Probe coverage today: en-IN, hi-IN, and te-IN. Latency does not affect Elo.
Integrity and transparency
• Every vote is attributed and timestamped; rankings reflect the full recorded history.
• Scripts, models, reference recordings, and generated clips are curated and versioned by the VoxClash team.
• Elo reflects human preference. TTFA reflects technical latency. Issue tags are qualitative diagnostics. We keep these separate so each number means one thing.

This doesn’t have to end here

Accuracy is Intelligence