January 2023 Summaries
8 posts from Speechmatics
Filter
Month:
Year:
Post Summaries
Back to Blog
Speechmatics' speech recognition technology is being used by Udemy to help aid and assist those teaching the next generation, with a focus on inclusivity and accessibility. The company's best-in-class speech-to-text tools are helping to meet the growing demand for accessible learning platforms, driven in part by tighter rules and regulations around accessibility. With 75,000 instructors in over 75 languages, Udemy is working to make its videos available in all languages, while ensuring that captions are accurate and comply with legislations such as section 508 and WCAG. The company's goal is to make learning accessible globally for everyone, and to achieve this it aims to provide high-quality captions at a price point that makes learning affordable. With the rise of online learning, captioning has become an essential aspect of providing trust and accessibility in educational content, with Udemy working closely with its instructors to ensure that captions meet the highest standards of accuracy and compliance.
Jan 31, 2023
1,204 words in the original blog post.
At Speechmatics, we aim to unlock human potential and increase inclusivity in speech recognition engines by integrating self-supervised learning into our technology. This advancement is made possible through the research and development of our machine learning team. The UK government's announcement of increased R&D budget credit was followed by proposed tax cuts that could block startups from investing and growing, which would negatively impact Speechmatics' business and investment in its own technology. Despite this, a letter from SME business leaders may yet see the cuts overturned, and universities are also concerned about the long-term damage R&D cuts can bring to their workforce and innovation ecosystem. The company believes that the government's contribution to R&D funding should not decrease further, as it is currently around £3.1bn, which contributes 22% of total R&D funding in the UK.
Jan 24, 2023
666 words in the original blog post.
Red Bee Media is a company that provides media services to broadcasters and content owners, including supply, enrichment, and show services such as subtitles, closed captions, audio description, and sign language translation. They work with Speechmatics' speech recognition technology to deliver high-quality accessibility for streaming services and the broadcast television industry. The partnership has enabled Red Bee Media to provide over 200 hours of live captioning a month for BT Sport, making it one of the most accessible sporting channels in the world. With advancements in automatic speech recognition, Red Bee Media is improving their automated captioning solution product to increase efficiency and accessibility across a wider range of content. They are also exploring future use cases such as music and quizzes, where speech recognition can be particularly challenging. The goal is to strike a balance between automation and human expertise to provide the best possible experience for audiences with sensory impairments.
Jan 19, 2023
1,446 words in the original blog post.
When it comes to speech-to-text systems, there's a significant abundance of data available online, especially in common languages, but under-resourced languages face a major challenge due to limited data availability. To address this issue, companies like Speechmatics turn to existing datasets such as Common Voice and OSCAR to fill the gaps and rapidly deploy new language support. The Common Voice project allows users to contribute labeled data, while OSCAR provides a multilingual corpus created from another open-source project, Common Crawl. By providing these datasets, both projects help bring inclusivity and equity to speech-to-text systems, enabling companies like Speechmatics to improve the accuracy of their models and support more languages. This collaboration enables rapid language deployment, reduces bias in content availability, and promotes equality for users worldwide.
Jan 17, 2023
723 words in the original blog post.
The development of speech technology has made significant progress in recent years, with AI-led speech-to-text becoming the dominant choice for transcription at scale. However, despite its impressive capabilities, speech recognition can be a barrier to accessibility, particularly for non-native speakers and individuals from diverse linguistic backgrounds. Research has shown that factors such as language, accent, race, gender, and age significantly impact the accuracy of speech recognition, leading to unequal access to digital technologies and a digital divide. The lack of representation of underrepresented languages in academia and industry exacerbates this issue, with modern deep learning systems relying heavily on data to achieve accuracy. To address these challenges, companies like Speechmatics are working towards making speech technology more inclusive and accessible, supporting 48 languages and aiming to expand coverage and improve technology to help interact with the digital world fairly for all people, regardless of language or background.
Jan 12, 2023
740 words in the original blog post.
Speechmatics aims to understand every voice with its accurate transcription engine and collaborates with partners like Prosodica to provide insights from customer conversations. The partnership has been in place since 2018, providing high-quality transcripts that support machine learning classifiers used by Prosodica for optimizing business operations. This collaboration enables the generation of satisfaction scores, assessment of soft skills performance, and predictions about risk or opportunity presented within a conversation. High-accuracy transcription is crucial as it affects machine learning accuracy, with more complex predictions being more sensitive to errors. The importance of diverse voice models has become apparent in recent years, especially for multinational corporations dealing with various voices and accents. In the future, speech analytics technology will continue to evolve, with contact centers adopting remote-ready solutions that support both human and automated interactions.
Jan 10, 2023
1,018 words in the original blog post.
In order to increase the return on investment in transcription tools, businesses need accurate and efficient speech-to-text solutions that can differentiate themselves from competitors and cater to a global market with diverse languages and accents. AI-based solutions are now the only viable future for transcription tools, offering significant cost savings compared to human transcription services. Speechmatics provides comprehensive speech-to-text solutions with features such as real-time transcription, speaker diarization, entity formatting, custom dictionary, profanity tagging, confidence scoring, and more. By choosing Speechmatics, businesses can improve their transcription tools, stand out in the market, and deliver on evolving customer expectations with accurate and efficient speech-to-text technology.
Jan 05, 2023
610 words in the original blog post.
Word Error Rate (WER) is a widely used metric to evaluate Automatic Speech Recognition (ASR) systems, but it has several flaws, including giving unequal importance to different types of mistakes and being sensitive to formatting differences. The calculation of WER involves aligning the reference and recognized transcript using Levenshtein distance, then counting substitutions, insertions, and deletions. However, the metric can be misaligned with reality due to issues such as punctuation, capitalization, and variations in writing styles. These limitations make it challenging for developers to use WER as a reliable evaluation tool, especially when comparing across different vendors. In recent work, Speechmatics is exploring novel metrics that align with human judgment by harnessing the power of large language models, offering a more meaningful way to evaluate ASR systems.
Jan 02, 2023
840 words in the original blog post.