Top-left quadrant shows models with fast generation and high ratings
DesignArenaby
Models
Methodology
In the Text-to-Speech Arena, every battle is a blind, head-to-head comparison: real users listen to two anonymous clips of the same prompt and vote for the one they prefer. Voice delivery can optionally be steered in a number of ways.
•Optional voice description: Steer the delivery (e.g. "warm, calm narrator with a British accent"). One description applies to both models.
•Gender + accent routing: For models that select from a fixed voice set, we route to a matching voice (gender first, accent second), falling back to a random pick when the description gives no signal.
•Free-form style steering: For models that support it, we also pass the full description through for age, emotion, and pacing.
•Fair matchmaking by capability: Descriptions requiring richer control only pit models that can honor them against each other, so no model is judged on a feature it lacks.
•Language-aware matching: The prompt's language is detected and models are only entered if they support it.