The Audio Realism Benchmark measures one thing: can a text-to-speech system sound indistinguishable from a real person in natural, everyday speech? Rather than traditional mean opinion scores, we ask the most direct question — of two voices reading the same line, which one sounds more human?
Expert native-speaker listeners hear two clips blind, in randomized left/right order, and pick the more human one. Pairwise votes are fit with the Bradley-Terry model to produce an Elo rating for each system. A real human recording of every prompt enters the same pool as an anchor, so beyond the headline ranking we can measure how often each model is chosen over an actual person.
Prompts are drawn from established spoken-language datasets and span three speech registers — phone agents (transactional service calls), conversations (podcasts and interviews), and explainers (teaching or clarifying an idea). No transcript is used verbatim: entities and topics are swapped so nothing is searchable back to its source, while disfluencies and natural speech patterns are preserved. Within each register, prompts are stratified by lexical difficulty so models are tested on both smooth and harder passages.
For each model we pin one female and one male voice — the provider's most naturally conversational American-English option — and every matchup is same-gender so accent and voice choice are never the variable under test.
Read the full methodology →