Dev News Daily ENDE

Hugging Face launches an open TTS leaderboard that ranks by metrics, not votes

Hugging Face has launched the Open TTS Leaderboard, a ranking of text-to-speech models built on objective metrics instead of listener votes. The motivation is volume: the Hub now holds more than 8,000 TTS models, and arena-style leaderboards that rank by human preference cannot keep up.

The authors also argue that arenas under-represent open models. As of 30 September, they count 16 open-weights models among 92 on Artificial Analysis, with a similar skew on Voice Arena. Their explanation is practical: adding an API model takes little more than a key, while an open model has to be hosted by the arena operator. They add that arenas cannot guarantee that voters apply the same idea of "better" over time.

The new leaderboard measures three things. Intelligibility is the word or character error rate between the prompt and a transcript of the generated audio, produced by Qwen3 ASR. Speed is the inverse real-time factor for batched inference on an H200 GPU, plus time to first audio at batch size one on GPU and CPU. For voice cloning, speaker similarity is the cosine similarity between WavLM speaker embeddings of the output and the reference clip. Evaluating a model takes a couple of hours instead of the weeks needed to collect votes.

The default view ranks models by average error rate on the English splits of Seed TTS Eval and CV3 Eval. There, hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead. Other languages can be toggled; k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are named as strong multilingual models.

The authors are explicit about the limits. The metrics are proxies: none of them measures naturalness, expressiveness or listener preference, and the leaderboard is meant to complement human ranking, not replace it.

Hugging Face launches an open TTS leaderboard that ranks by metrics, not votes
Hugging Face launches an open TTS leaderboard that ranks by metrics, not votes — Dev News Daily

Why it matters

A model that tops an arena may never have been tested on the language you ship. A reproducible, per-language error rate is a cheaper first filter, as long as nobody mistakes "fewest transcription errors" for "sounds best".