Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

A Real-World Dataset for Noise-Robust Speech AI: 100+ Timestamped Noise Types Across 58 Indian Languages

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Suryansh Shukla, PavanKumarJ, Sujith Pulikodan, Agneedh Basu, Pranav D Bhat, and Prasanta Kumar Ghosh
Word Count
1,184
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

ARTPARK-IISc’s Vaani Noise Event Timestamp Dataset adds a human-annotated noise layer to Project Vaani’s multilingual Indian speech collection, providing 106,892 real-world noise events with millisecond-level start and end timestamps across more than 122 hours of audio, 58 languages, 38,541 speakers, and 30 states. Recorded naturally on mobile devices in field conditions rather than created through synthetic noise mixing, the dataset captures overlapping speech and everyday sounds such as traffic, animals, children, music, alarms, appliances, and non-speech human noises. It includes roughly 22 hours of multi-annotator-verified data and about 100 hours of unverified data, supported by a multi-stage quality-control process. The dataset is positioned as addressing a gap left by existing speech and noise resources, which generally lack the combination of natural co-occurring noise, precise temporal labels, and broad Indic-language coverage. Its intended uses include improving noise-robust speech recognition, sound event detection, speech enhancement, and research for low-resource languages, with the broader aim of making voice AI more reliable in everyday Indian environments.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Voice AI 6 324 41 16 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.