July 2023 Summaries
7 posts from Speechmatics
Filter
Month:
Year:
Post Summaries
Back to Blog
Trevor Back, a seasoned machine learning expert, has joined Speechmatics as Chief Product Officer, bringing over a decade's experience and a track record of growth and innovation to the role. He will lead product strategy and execution, leveraging his ability to translate AI research into real-world production software. Usman Gulfaraz, former Vice President of Sales EMEA/APAC at Tessian, has been appointed as Chief Revenue Officer, bringing substantial growth expertise and industry knowledge to accelerate revenue generation and global expansion efforts. The appointments come at a pivotal moment for Speechmatics, which has recently closed a $62 million funding round and launched several new products and services, including Ursa, real-time translation and transcription, and summarization. The additions are seen as key to driving innovation globally and setting the market standard for speech-to-text and translation offerings.
Jul 28, 2023
663 words in the original blog post.
The internet has seen numerous captioning fails, often due to human errors or inaccuracies in automatic speech recognition (ASR) technology. These mistakes can harm a brand's reputation, limit accessibility and inclusion, and reduce viewer engagement. However, investing in accurate captions is crucial for promoting digital inclusion and enhancing user experience. Technology, such as real-time ASR and machine learning, can help mitigate errors and ensure timely and synchronized captions, thereby elevating brands to the next level of accessibility and inclusivity.
Jul 25, 2023
1,231 words in the original blog post.
Speechmatics has launched its Summarization feature, a comprehensive suite of speech understanding features for customers, which utilizes abstractive summarization to analyze input, extract key points, and produce a summary that captures the essence of the content. The feature offers flexibility in output type, length, and format, catering to various use cases such as contact centers, virtual meetings, media, and podcasts. Large Language Models (LLMs) power the Summarization tool, enabling it to generate coherent summaries based on contextual understanding and creative text generation capabilities. Despite limitations, Speechmatics' Summarization can handle files of any duration, making it a valuable tool for improving productivity and collaboration in content creation and analysis.
Jul 24, 2023
1,133 words in the original blog post.
Large language models are AI models trained on vast amounts of text data, learning patterns and generating human-like text, answering questions, summarizing text, and more. These models can understand multiple phrases and topics, helping businesses explore abundant interactions with transcripts to gather valuable insights. The Speechmatics team has been exploring summarization across multiple transcripts, gathering key takeaways, action items, company blockers, and broader themes, and is discussing hosting open-source language model on premises to solve security issues. LLMs offer numerous opportunities across industries, including media, EdTech, and calls, with use cases such as content creation, advanced archiving, robust monitoring, extracting course content from audio, faster note-taking/revision, summarizing outcomes, extracting resolutions, highlighting outstanding issues, and more. The future of LLMs holds immense potential for driving substantial gains in businesses, with Speechmatics well-positioned to support customers in leveraging these capabilities.
Jul 24, 2023
804 words in the original blog post.
The distinction between closed captioning and open captioning lies in their display options and flexibility. Closed captions are optional and can be enabled or disabled, whereas open captions are permanently displayed on the screen and cannot be turned off. Open captions ensure universal accessibility but may impact visual aesthetics, while closed captions offer customization options but require manual setup. Speech-to-text technology has streamlined the captioning process, making it more efficient, scalable, cost-effective, and inclusive by automatically generating captions for large volumes of content, enabling real-time captioning, reducing dependency on manual labor, and addressing common challenges such as accuracy and latency issues.
Jul 20, 2023
1,253 words in the original blog post.
YouTube's auto-captioning system is notoriously unreliable, with accuracy rates ranging from 60-70%, which means that up to 30% of captions may be incorrect. This has significant implications for content creators who rely on accurate captions, particularly in educational and potentially life-saving videos. To address this issue, Speechmatics has developed an AI-powered speech-to-text engine that demonstrates significantly higher accuracy rates, with some tests showing levels above 90%. The company's self-supervised learning approach allows it to improve its engine by incorporating a vast amount of unlabeled data, which helps bridge the gap between well-curated and everyday speech. As captioning becomes increasingly important, the market is shifting towards innovation, prioritizing accuracy, and making captions a necessity rather than an add-on.
Jul 08, 2023
812 words in the original blog post.
Hyperscalers, such as Microsoft, offer convenient ASR solutions with generic models that have limitations for Contact Center as a Service (CCaaS) vendors. In contrast, specialist ASR providers like Speechmatics deliver high accuracy for every speaker, supporting multiple languages and downstream features. CCaaS vendors require accurate ASR to provide a superior customer experience and grow in non-English markets. Hyperscalers may limit competition, while specialized ASR providers offer flexibility, accuracy, and support. Partnering with Speechmatics allows for customized solutions and continuous innovation in contact center operations. The company's ASR solution has been developed using Ursa generation models and delivers a relative accuracy lead of up to 30% compared to Amazon, Google, IBM, and Microsoft. Speechmatics' highly accurate ASR improves the quality of downstream tasks, enabling CCaaS providers to boost performance, uncover valuable customer insights, enhance quality monitoring, and optimize total cost of ownership for their infrastructure. The company also offers self-supervised learning, which enables it to train on much more unlabeled data and perform better on audio with background noise. Furthermore, Speechmatics values offering exceptional experiences to its customers by providing a highly versatile ASR model that can be tailored to solve specific requirements and pain points.
Jul 04, 2023
1,921 words in the original blog post.