Blind Voice Comparison
Hear the same script from two anonymous models, then vote for the one that sounds more natural. Names are revealed only after you vote.
Each round picks a script at random from our test set — weighted so every domain gets fair coverage over time.
🧪 Custom text — counts toward Overall only, not a specific domain
Listen to both, then cast your vote ↓
Which voice sounds more natural?
Leaderboard
Methodology
Evaluation Setup
All models receive the same input text for each comparison. For every round, two models are randomly selected and generate speech in parallel.
Evaluators listen to anonymized audio samples labelled only as Model A and Model B to ensure a fair comparison. After listening, they vote for the sample they believe has better overall speech quality. The model identities are revealed only after the vote is submitted.
Rankings come from a Bradley-Terry model — the same statistical approach used by Chatbot Arena and Voice Arena — refit from every recorded vote each time a new one comes in, rather than nudged step-by-step after each match. Ties count as half a win for each model. Ratings are displayed on the familiar Elo scale, centered at 1200.
Questions
Credits
This arena's blind side-by-side format was inspired by TTS Arena on Hugging Face. Our preset category taxonomy and Bradley-Terry ranking methodology follow the approach documented by Voice Arena.