Home / Companies / Speechmatics / Blog / Post Details
Content Deep Dive

How to Rapidly Train New Languages Using Common Voice and OSCAR

Blog post from Speechmatics

Post Details
Company
Date Published
Author
Steve Kingsley
Word Count
723
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

When it comes to speech-to-text systems, there's a significant abundance of data available online, especially in common languages, but under-resourced languages face a major challenge due to limited data availability. To address this issue, companies like Speechmatics turn to existing datasets such as Common Voice and OSCAR to fill the gaps and rapidly deploy new language support. The Common Voice project allows users to contribute labeled data, while OSCAR provides a multilingual corpus created from another open-source project, Common Crawl. By providing these datasets, both projects help bring inclusivity and equity to speech-to-text systems, enabling companies like Speechmatics to improve the accuracy of their models and support more languages. This collaboration enables rapid language deployment, reduces bias in content availability, and promotes equality for users worldwide.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 297 62 31 +7%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.