Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Language Identification for 42 Indian Languages: Opening New Frontiers in Indic Speech AI

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Suryansh Shukla, Agneedh Basu, Sujith Pulikodan, Pranav D Bhat, and PavanKumarJ
Word Count
1,370
Company Posts That Month
14
Language
-
Hacker News Points
-
Post removed?
No
Summary

ARTPARK-IISc has released Vaani-LID_v0, an open MIT-licensed spoken language identification model for 42 Indian languages spanning Indo-Aryan, Dravidian, Sino-Tibetan, and English, trained using the Vaani corpus. The study used balanced 10-hour samples per language with speaker- and district-disjoint training, validation, and test sets to reduce the risk of models recognizing speakers rather than languages, and evaluated Whisper and a Vaani-pretrained FastConformer encoder on in-domain and out-of-domain benchmarks. A frozen FastConformer pretrained on geographically and demographically diverse Indic speech substantially outperformed fine-tuned Whisper on the external FLEURS and Kathbath datasets, while fine-tuned Whisper performed better on the in-domain Vaani test set, suggesting that broad regional pretraining can improve generalization but further fine-tuning may reduce it. Hierarchical softmax, which models linguistic family and subfamily relationships before individual languages, improved results across encoders and datasets compared with standard classification objectives. Performance varied sharply by language family, with Sino-Tibetan languages achieving the highest accuracy and closely related Central Indo-Aryan varieties proving most difficult, particularly Hindi versus Urdu and Sadri, Chhattisgarhi, and Surgujia. The authors have also made the 31,255-hour Vaani dataset and related multilingual ASR resources publicly available, inviting further work on underperforming low-resource varieties.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 2 103 37 26 -89%
Voice AI 2 1,179 83 25 -73%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.