YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research
Blog post from Hugging Face
YODAS v3 is a newly released 1.1 million-hour multilingual speech dataset intended to expand open research access to real-world voice data at a scale comparable to major proprietary collections. Developed by the ESPnet community and licensed under CC BY 3.0, it triples the size of the original YODAS dataset while covering more than 100 languages, including 34 with over 1,000 hours of audio. The release provides 48 kHz audio, with most recordings offering effective 32 kHz or higher frequency content, and more than 70% containing genuinely distinct multichannel audio. Each example includes transcripts with word- and utterance-level timestamps, detected language labels, and, for more than half of multilingual material, timestamped English translations. Packaged in 60 TB of OPUS files, YODAS v3 is designed to support multilingual speech recognition, speech synthesis, spatial audio generation, speech translation, alignment, audio codec development, and speech enhancement, building on earlier YODAS versions that have been widely downloaded and used in datasets, benchmarks, and foundation models.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.