Text-to-Speech

DesignArenaby Intelligence

Loading chart data...

The Leaderboard

Methodology

In the Text-to-Speech Arena, every battle is a blind, head-to-head comparison: real users listen to two anonymous clips of the same prompt and vote for the one they prefer. Voice delivery can optionally be steered in a number of ways.

  • Optional voice description: Steer the delivery (e.g. "warm, calm narrator with a British accent"). One description applies to both models.
  • Gender + accent routing: For models that select from a fixed voice set, we route to a matching voice (gender first, accent second), falling back to a random pick when the description gives no signal.
  • Free-form style steering: For models that support it, we also pass the full description through for age, emotion, and pacing.
  • Fair matchmaking by capability: Descriptions requiring richer control only pit models that can honor them against each other, so no model is judged on a feature it lacks.
  • Language-aware matching: The prompt's language is detected and models are only entered if they support it.
Design Arena | Leaderboards