//  VoXClash

Blind pairwise evaluation across languages.

Model-vs-model Elo rankings, human reference anchors, and per-language quality metrics for state-of-the-art TTS systems.
Hindi
English
Telugu
Model leaderboard
Hindi · Elo from model-vs-model votes only
Head-to-Head win-rates
Row model's win % against column model
Compare-to-human voting is coming soon
Model-vs-model rankings are live above. Expert human recordings will be integrated as benchmark baselines soon.
Latency
P50 is the usual wait. P25–P75 is how jumpy that wait is. Min–max is best vs worst. Silence is hush at the start of the clip.
Issue tags by model
Tagged issues as a share of each model's model-vs-model matchups
Model leaderboard
English (IN) · Elo from model-vs-model votes only
Head-to-Head win-rates
Row model's win % against column model
Compare-to-human voting is coming soon
Model-vs-model rankings are live above. Expert human recordings will be integrated as benchmark baselines soon.
Latency
P50 is the usual wait. P25–P75 is how jumpy that wait is. Min–max is best vs worst. Silence is hush at the start of the clip.
Issue tags by model
Tagged issues as a share of each model's model-vs-model matchups
Model leaderboard
Telugu · Elo from model-vs-model votes only
Head-to-Head win-rates
Row model's win % against column model
Compare-to-human voting is coming soon
Model-vs-model rankings are live above. Expert human recordings will be integrated as benchmark baselines soon.
Latency
P50 is the usual wait. P25–P75 is how jumpy that wait is. Min–max is best vs worst. Silence is hush at the start of the clip.
Issue tags by model
Tagged issues as a share of each model's model-vs-model matchups

Frequently Asked Questions

  • Rankings come from blind listening tests: people hear two clips of the same script and pick which sounds better. Scores are Elo ratings from those pairwise choices, and not automated text or speech metrics.

  • Higher Elo means the model wins more head-to-head comparisons. The bracket is a 95% uncertainty range: with current votes, the score could reasonably sit within that band.

  • Latency measures how quickly audio starts (time-to-first-audio). It does not affect quality rank, a model can be fast but sound worse, or slow but sound better.

This doesn’t have to end here

Accuracy is Intelligence