Language Identification for 42 Indian Languages: Opening New Frontiers in Indic Speech AI
Blog post from Hugging Face
ARTPARK-IISc has released Vaani-LID_v0, an open MIT-licensed spoken language identification model for 42 Indian languages spanning Indo-Aryan, Dravidian, Sino-Tibetan, and English, trained using the Vaani corpus. The study used balanced 10-hour samples per language with speaker- and district-disjoint training, validation, and test sets to reduce the risk of models recognizing speakers rather than languages, and evaluated Whisper and a Vaani-pretrained FastConformer encoder on in-domain and out-of-domain benchmarks. A frozen FastConformer pretrained on geographically and demographically diverse Indic speech substantially outperformed fine-tuned Whisper on the external FLEURS and Kathbath datasets, while fine-tuned Whisper performed better on the in-domain Vaani test set, suggesting that broad regional pretraining can improve generalization but further fine-tuning may reduce it. Hierarchical softmax, which models linguistic family and subfamily relationships before individual languages, improved results across encoders and datasets compared with standard classification objectives. Performance varied sharply by language family, with Sino-Tibetan languages achieving the highest accuracy and closely related Central Indo-Aryan varieties proving most difficult, particularly Hindi versus Urdu and Sadri, Chhattisgarhi, and Surgujia. The authors have also made the 31,255-hour Vaani dataset and related multilingual ASR resources publicly available, inviting further work on underperforming low-resource varieties.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 2 | 103 | 37 | 26 | -89% |
| Voice AI | 2 | 1,179 | 83 | 25 | -73% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.