AiPhreaks ← Back to News Feed

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

By Jakub Antkiewicz

•

2026-10-01T15:07:54Z

New Leaderboard Aims to Standardize Open-Source TTS Evaluation

A new Open TTS Leaderboard has been launched to address the fragmented and unstandardized evaluation landscape for text-to-speech models. With over 8,000 TTS models available on the Hugging Face Hub, the pace of open-source releases has far exceeded the capacity of traditional human-preference-based systems like TTS Arena v2 and Voice Arena. These arena-style leaderboards, while valuable, struggle to scale and tend to underrepresent open-weight models, which require hosting and serving by the arena operator. The new leaderboard provides an objective, automated framework intended to complement, not replace, human feedback.

Objective Metrics for Speed, Intelligibility, and Similarity

The Open TTS Leaderboard evaluates models using a suite of objective metrics, allowing for rapid and reproducible comparisons that take hours instead of weeks. The evaluations focus on several key aspects of performance, measured on standardized datasets like Seed TTS Eval and CV3 Eval. Top-performing English models currently include hexgrad/Kokoro-82M and fishaudio/s2-pro, while k2-fsa/OmniVoice shows strong multilingual capabilities.

  • Intelligibility: Word Error Rate (WER) and Character Error Rate (CER) are calculated by comparing the generated audio's transcript against the original prompt, using the Qwen3 ASR model.
  • Speed: Performance is measured by inverse real-time factor (RTFx) for batched inference and time-to-first-audio (TTFA) for streaming latency on an H200 GPU and CPU.
  • Speaker Similarity: For voice cloning, cosine similarity (SIM) between WavLM speaker embeddings of the reference and generated audio is computed to measure identity preservation.

Shaping a More Transparent AI Ecosystem

By focusing on open-source and multilingual models, the initiative aims to surface high-quality systems that might otherwise be overlooked by commercial-centric evaluation platforms. The leaderboard provides dedicated views for multilingual performance, voice cloning, and streaming capabilities, which are critical for interactive applications. While objective metrics for intelligibility and similarity do not directly measure subjective qualities like naturalness or expressiveness, they provide a crucial, scalable proxy. The platform also includes a “Listen” tab where the community can compare outputs and vote, creating a valuable feedback loop to inform future development and evaluation standards.

The launch of the Open TTS Leaderboard signals a necessary maturation in AI evaluation, where scalable, objective benchmarks become critical infrastructure for navigating the flood of open-source models. This shifts the focus from purely subjective preference battles toward reproducible, multi-faceted performance analysis that accelerates both development and adoption.
End of Transmission
Scan All Nodes Access Archive