Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research

Blog post from Hugging Face

Post Details
Company
Date Published
Author
William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, and Shinji Watanabe
Word Count
1,307
Company Posts That Month
71
Language
-
Hacker News Points
-
Post removed?
No
Summary

YODAS v3 is a newly released 1.1 million-hour multilingual speech dataset intended to expand open research access to real-world voice data at a scale comparable to major proprietary collections. Developed by the ESPnet community and licensed under CC BY 3.0, it triples the size of the original YODAS dataset while covering more than 100 languages, including 34 with over 1,000 hours of audio. The release provides 48 kHz audio, with most recordings offering effective 32 kHz or higher frequency content, and more than 70% containing genuinely distinct multichannel audio. Each example includes transcripts with word- and utterance-level timestamps, detected language labels, and, for more than half of multilingual material, timestamped English translations. Packaged in 60 TB of OPUS files, YODAS v3 is designed to support multilingual speech recognition, speech synthesis, spatial audio generation, speech translation, alignment, audio codec development, and speech enhancement, building on earlier YODAS versions that have been widely downloaded and used in datasets, benchmarks, and foundation models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Voice AI 3 324 41 16 -89%
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.