The Open ASR Leaderboard Adds Its First Global South Language
Blog post from Hugging Face
Voice Arena and Hugging Face have added Monsoon en-IN and Monsoon hi-IN to the Open ASR Leaderboard, introducing Indian English and Hindi evaluation sets designed to reveal speech-recognition performance differences that aggregate word error rates can obscure. The public and private, speaker-disjoint splits include 4,888 speakers across hundreds of Indian districts, varied devices, acoustic settings, demographic backgrounds, and conversational speech styles, with extensive per-segment metadata supporting analysis by region, age, gender, education, occupation, and handset. The datasets prioritize breadth of speaker and geographic representation over long recordings from a small number of contributors, while using screening, recording checks, quality controls, and multi-stage native-linguist transcription to improve reliability. For Hindi, the benchmark uses transcript lattices and Orthographically-Informed Word Error Rate to accommodate legitimate spelling and code-mixing variations that conventional single-reference WER can penalize unfairly. An example analysis of Indian English shows that models with nearly identical overall scores can differ substantially by speakers’ regions, illustrating how the new sets aim to make demographic and linguistic variation visible in widely used ASR evaluations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 1 | 2,814 | 261 | 53 | -37% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.