How to Rapidly Train New Languages Using Common Voice and OSCAR
Blog post from Speechmatics
When it comes to speech-to-text systems, there's a significant abundance of data available online, especially in common languages, but under-resourced languages face a major challenge due to limited data availability. To address this issue, companies like Speechmatics turn to existing datasets such as Common Voice and OSCAR to fill the gaps and rapidly deploy new language support. The Common Voice project allows users to contribute labeled data, while OSCAR provides a multilingual corpus created from another open-source project, Common Crawl. By providing these datasets, both projects help bring inclusivity and equity to speech-to-text systems, enabling companies like Speechmatics to improve the accuracy of their models and support more languages. This collaboration enables rapid language deployment, reduces bias in content availability, and promotes equality for users worldwide.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 1 | 297 | 62 | 31 | +7% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.