Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

SraVaani: How Vision Helps Hearing

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Agneedh Basu, Sujith Pulikodan, PavanKumarJ, Pranav D Bhat, and Suryansh Shukla
Word Count
1,648
Company Posts That Month
74
Language
-
Hacker News Points
-
Post removed?
No
Summary

SraVaani is a multilingual Indian automatic speech recognition model from ARTPARK-IISc that supports 65 languages and dialects and is built using the Vaani corpus, which contains more than 31,000 hours of speech across 105 languages but relatively limited transcription coverage. Its central approach adds an audio-image alignment stage between self-supervised audio pretraining and supervised ASR fine-tuning, exploiting the corpus’s picture-prompt collection method to teach a FastConformer speech encoder to associate spoken descriptions with related image embeddings without requiring new transcripts. The model uses a 17-layer FastConformer encoder, frozen SigLIP2 vision representations during alignment, and a hybrid CTC-TDT decoder for final speech recognition training on 30,565 hours of labeled speech from 18 public datasets. On a held-out Vaani evaluation spanning 48 languages, the alignment method reduced word error rate from 28.09% to 27.40%, with reported improvements in 37 languages and particularly strong coverage for lower-resource tribal and dialect languages where comparison systems often produced no output. SraVaani-1.0 is released through Hugging Face in NeMo format, alongside training and evaluation resources for adaptation to other speech datasets.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 6 516 143 56 -47%
Vector Search 1 2,312 357 123 +3%
Voice AI 1 2,814 261 53 -37%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.