The Open ASR Leaderboard Adds Its First Global South Language
By Jakub Antkiewicz
•2026-08-29T13:38:04Z
Open ASR Leaderboard Adds Hindi and Indian English to Tackle Model Bias
Hugging Face and Voice Arena have launched two new evaluation sets for Hindi and Indian English on the Open ASR Leaderboard, marking the first inclusion of a Global South language. The move directly addresses a critical weakness in automated speech recognition development: benchmarks that ignore demographic and regional diversity often lead to models that perform poorly for specific populations. By introducing the Monsoon datasets, the leaderboard now provides a public tool to measure and mitigate the known disparities where commercial ASR systems can have error rates twice as high for underrepresented groups.
The Monsoon collection is engineered for variance, not just volume. While modest in total hours, it is extensive in its speaker diversity, comprising 4,888 unique speakers across its four public and private splits. This design prevents models from overfitting to a small number of voices. The data was sourced from unscripted conversations recorded on contributors' own devices across hundreds of Indian districts, capturing authentic acoustic conditions and speech patterns. Each segment is enriched with 12 metadata attributes, allowing for granular performance analysis.
- Languages: Hindi (hi-IN) and Indian English (en-IN)
- Total Speakers: 4,888 speaker-disjoint participants
- Design Axes: Varies by geography, age, gender, devices, acoustic environments, and more
- Transcription: Standard text for English; a transcript lattice for Hindi to accept multiple correct spellings
- Data Source: Spontaneous, dual-channel conversations on everyday topics
The impact of this launch extends beyond the technical specifications. By making fine-grained, demographically-aware evaluation a public standard, the Monsoon benchmark incentivizes the AI industry to build more equitable ASR systems. Previously, this level of analysis was confined to private, internal datasets. Now, any developer competing on the Open ASR Leaderboard must confront how their models perform across the diverse accents, dialects, and device profiles prevalent in the Indian subcontinent. This shift forces a focus from optimizing a single Word Error Rate (WER) to building robust models for one of the world's largest and most varied linguistic markets.
The addition of the Monsoon dataset to the Open ASR Leaderboard is a market-correcting mechanism. By institutionalizing the measurement of performance across diverse demographics and real-world conditions, it moves the goalposts for 'state-of-the-art' from raw WER to equitable, commercially-viable ASR for over half a billion new users.