Home / Companies / Deepgram / Blog / December 2021

December 2021 Summaries

24 posts from Deepgram

Filter
Month: Year:
Post Summaries Back to Blog
In this article, the author shares their experience of building a cross-platform NuGet package for the .NET ecosystem. They discuss the requirements and challenges faced while creating an SDK that supports various frameworks and platforms. The author also provides insights into using GitHub Actions for continuous deployment (CD) to NuGet.org when a new version is released. Finally, they announce the release of the Deepgram .NET SDK and invite feedback and contributions from the community.
Dec 23, 2021 1,272 words in the original blog post.
Multichannel audio and diarization are two features that can be used for speech recognition tasks. Multichannel audio separates different people's voices into individual channels, making it easier to focus on one speaker when reviewing the audio file. Diarization, on the other hand, separates audio by the person speaking, whether they are on a different channel or not. Deepgram offers both multichannel and diarization features that can be used individually or combined depending on the specific use case.
Dec 20, 2021 2,017 words in the original blog post.
This tutorial guides users through building a voice-powered song search using Deepgram and the Genius Song Lyrics API. The project involves setting up a server, accessing and sending audio data from the user's microphone to be processed by Deepgram, and displaying results on a website. Users will learn how to stream microphone data to Deepgram via a server, protecting their API Key from exposure. The final product is a voice-activated song search engine that can correctly guess spoken or sung lyrics.
Dec 16, 2021 1,445 words in the original blog post.
Deepgram has been named a High Performer in Voice Recognition Software by G2 for the second consecutive quarter, rising to the number two spot based on user reviews. The company received high ratings from users, with 91% believing it's heading in the right direction and 89% likely to recommend Deepgram to others. Users praised Deepgram's API coverage, accuracy, speed, and ease of use.
Dec 14, 2021 246 words in the original blog post.
The MediaStream API is a crucial tool for web developers when working with audio and video inputs. This post provides an overview of the basics of the MediaStream API, including getting started, properties, methods, and events. Key points include understanding how to gain access to user's audio/video devices using getUserMedia method, the concept of a 'stream' consisting of one or more 'tracks', various properties like active and id, methods such as addTrack, getTracks, removeTrack, and events like onaddtrack and onremovetrack. Understanding these concepts can help developers effectively utilize the MediaStream API in their applications.
Dec 13, 2021 990 words in the original blog post.
The evolution of conversational AI in the car and beyond has seen significant advancements since its inception in 2004 with Honda collaborating with IBM to launch a voice navigation system. Over time, natural language understanding improved, and companies like Ford launched their own voice assistants. In 2013, Apple changed the game by integrating Siri into cars through CarPlay, followed by Android Auto in 2014. Amazon's Alexa entered the automotive space in 2017 with an embedded system for Ford vehicles, and later collaborated with other companies to create Voice Interoperability. In 2019, Amazon launched Alexa Auto SDK, which enabled more control over in-car systems. The future of conversational AI in cars is expected to include sentiment analysis, emotion recognition, and humanized technologies for a more personalized user experience.
Dec 09, 2021 2,365 words in the original blog post.
John Kelvie, CEO of Bespoken, discussed the importance of testing with voice experiences and conversational AI at Project Voice X. He emphasized that while AI is advanced, it's not autonomous or self-learning, and requires continuous training to improve performance. Kelvie presented a case study where they helped a customer improve their conversational AI system by using Azure speech instead of Lex for better recognition hints and error rates reduction. He also highlighted the need for trainable, testable, modular, and contextual conversational applications that can be easily improved and optimized over time. Kelvie's company, Bespoken, aims to support the entire life cycle of conversational AI systems through their platform, which includes analytics, management, orchestration, and continuous improvement.
Dec 09, 2021 2,703 words in the original blog post.
Braden Ream, CEO of Voiceflow, presented a session on "Sparking the Future of Conversation Design" at Project Voice X. He discussed the challenges of building a conversation design tool and how it can teach about conversation design itself. The talk covered the importance of conversation design in the automation industry, the limitations of existing tools, and the role of conversation designers. Ream also shared some customer stories that highlighted the need for better documentation and collaboration features in conversation design tools. He ended by discussing some problems the industry has yet to solve, such as intent management, content management, and persona management.
Dec 09, 2021 3,914 words in the original blog post.
Scott Sandland, CEO of Cyrano.AI, presented a keynote at Project Voice X on the importance of effective communication and how technology can be used to improve it. He discussed his experience working with high-functioning executives and drug-addicted teenagers, highlighting the similarities in their language patterns and thought processes. Sandland emphasized the need for more than just sentiment analysis in understanding a person's motivations and decision-making process. His company, Cyrano.AI, measures words and language in contextual ways to provide subjective understanding about an audience. The technology can be used in various applications such as personal assistants, customer service, and mental health support.
Dec 09, 2021 4,853 words in the original blog post.
Sonic branding is an art and science of creating strategic development and deployment of consistent authentic sound for enterprises. It involves research, testing, discovery, and working on the technology side to create a unified communication that strengthens brands and companies. The top enterprise companies know the importance of sonic branding in creating consistency across multiple touchpoints, differentiating themselves from competitors, and building emotional connections with customers. In recent years, there has been an increase in large-scale efforts across entire organizations, investment or partnership with voice, sound, and technology companies, and collaboration between various industries to leverage the power of sonic branding.
Dec 09, 2021 3,145 words in the original blog post.
Deepgram's API features Search and Keywords, which are designed for different scenarios. Search is a query that helps find if a word or phrase has been said in the audio being transcribed, while Keywords provide more information to improve transcription accuracy. The Search feature returns possible matches with confidence ratings, useful for compliance checks and eDiscovery. On the other hand, Keywords can boost specific words' recognition during transcription but require careful implementation and testing. Both features have their unique applications and can aid in better decision-making when analyzing audio transcriptions.
Dec 09, 2021 2,071 words in the original blog post.
Art Coombs, CEO of KomBea, discussed the future of AI in contact centers during a session at Project Voice X. He shared his experience working with Cisco founders Len Bosack and Sandy Lerner in 1982 when they introduced the concept of artificial intelligence (AI). Coombs emphasized that customers want quick, easy, and accurate service from both human agents and AI chatbots. He explained how technology has evolved over time to improve customer interactions, with a focus on cost reduction and improving CSAT scores. Coombs also highlighted the impact of COVID-19 on call center operations and the increasing role of AI in handling conversations. He believes that AI will consume technology in the future, making it crucial for companies to adapt their contact centers accordingly.
Dec 09, 2021 3,365 words in the original blog post.
Esteban Gorupicz, CEO of Atexto, discussed improving speech recognition accuracy through data processes at Project Voice X. He highlighted the importance of addressing issues like word error rate, bias, fairness, and language support in voice technologies. Atexto's platform enables companies to visualize, label, and collect speech training data faster, offering features such as diarization, custom vocabulary, redaction, punctuation, profanity filtering, and numeral formatting. The company also provides a crowdsourcing facility with over 1.5 million users for annotating speech and text, as well as ASR benchmark modules to measure word error rate and token error rate. Atexto's vision is to help companies adopt voice technologies and democratize them for all businesses in the future.
Dec 09, 2021 833 words in the original blog post.
Project Voice X 2021 took place in Destin, Florida, bringing together innovators and futurists discussing the advancements of voice technology. The event covered topics such as using voice interfaces for business, retail, and medical purposes, NLU-based multichannel communication, and analyzing voice personality traits. Presentations included "Building the Future of Voice" by Scott Stephenson (CEO, Deepgram), "The New Age of Voice Commerce" by Mike Zagorsek (COO, SoundHound), and "Voice in Healthcare Part II" by Henry O'Connell (CEO, Canary Speech).
Dec 09, 2021 451 words in the original blog post.
In the opening keynote of Project Voice X, Jeff Blankenberg discussed the future of voice technology and artificial intelligence (AI) assistants. He emphasized that people are at the center of all technological advancements, as they create products and services for human use. The speaker highlighted the importance of trust in AI systems, especially when it comes to handling sensitive information like financial transactions or personal health data. Blankenberg also touched upon various applications of AI in daily life, such as managing bills, scheduling appointments, refinancing houses, and enhancing social interactions. He argued that AI could help eliminate the need for subjective decision-making in certain areas, like medicine and finance, by providing more accurate diagnoses and predictions based on data analysis. The speaker also addressed the challenges of incorporating AI into everyday routines, emphasizing the importance of minimizing behavior changes to ensure widespread adoption. He provided examples of how AI could be integrated seamlessly into daily life, such as proactive notifications about groceries or smart home devices that anticipate user needs. In conclusion, Blankenberg expressed excitement about the potential for voice technology and AI assistants to solve problems without direct human input, freeing up time for users to pursue their interests and passions. He encouraged attendees to join the Alexa Design community on Slack to discuss these topics further.
Dec 09, 2021 5,671 words in the original blog post.
Henry O'Connell, CEO of Canary Speech, presented "Voice in Healthcare" at Project Voice X. He discussed how speech and language technology can be commercialized in the healthcare space to provide actionable information for diagnosis. Canary Speech has been issued eight patents and is working with over a dozen hospitals globally. The company's technology analyzes millions of datasets per month, focusing on depression, anxiety, stress, tiredness, cognitive functions, and Alzheimer's disease. Their tools are language agnostic and can be deployed in various languages. Canary Speech aims to streamline information for patient diagnosis, reduce readmissions, and provide immediate actionable guidance for healthcare providers.
Dec 09, 2021 2,561 words in the original blog post.
Blutag CEO Shilp Agarwal discussed the current state of voice commerce and its potential for growth in retail experiences. He highlighted that while there are significant numbers of people using smart speakers for shopping, many attempts to replicate full website functionality through voice have not been successful. Instead, focusing on tasks such as quickly reordering items or adding them to a cart has shown promise, with some customers seeing an increase of up to 35% in reorder frequency and 11-12% in shopping cart sizes. Agarwal emphasized the importance of providing value through voice technology and suggested that future opportunities may lie in areas like auto and restaurant ordering.
Dec 09, 2021 2,881 words in the original blog post.
Scott Stephenson, CEO of Deepgram, discussed the future of voice technology and its increasing adoption in call centers and meetings during his keynote at Project Voice X. He highlighted how automation and AI are transforming customer service by enabling real-time analysis of audio data, eliminating the need for surveys or text messages to gauge customer satisfaction. Stephenson emphasized that Deepgram's technology can handle various challenges in speech recognition, such as accents, dialects, background noise, and conversational styles. He also mentioned their start-up program offering up to $100,000 in speech-recognition credit for voice tech entrepreneurs.
Dec 09, 2021 3,855 words in the original blog post.
The New Age of Voice Commerce is a rapidly growing market with immense potential for businesses and consumers alike. With advancements in voice AI technology, companies are now able to extend their product experiences through voice using custom voice interfaces. This has led to the creation of new revenue streams and increased customer loyalty. As more organizations adopt voice strategies, it's crucial to focus on delivering value and creating monetization opportunities for both customers and businesses. By understanding user behavior and preferences, companies can create seamless, proactive purchasing options that enhance the overall customer experience while generating revenue. Ultimately, voice AI should generate revenue, not cost, and businesses should be asking themselves how much they can earn by moving forward in this technology rather than focusing on its costs.
Dec 09, 2021 4,157 words in the original blog post.
In a presentation at Project Voice X, Lenovo's Robert Daigle and Andi Huels discussed the integration of NLP (Natural Language Processing) on Edge devices. They emphasized the importance of leveraging technology partners to optimize solutions and improve pricing from hardware providers like NVIDIA and Intel. The speakers highlighted a use case involving an AI-powered voice assistant for drive-through ordering, which reduces friction for customers and speeds up the order process. They also mentioned Lenovo's support for internal data science teams through their innovation centers and conversational AI assistance for customer and employee support.
Dec 09, 2021 1,441 words in the original blog post.
Dr. Yared Alemu, CEO of TQIntelligence, discussed the impact of early life trauma on mental health and its subsequent effects on individuals transitioning into adulthood during Project Voice X. He highlighted that traumatic experiences can lead to a twenty-year reduction in life expectancy and various chronic diseases. Dr. Alemu emphasized the need for an objective way to measure and quickly assess the severity of mental health issues, as well as effective tracking methods to monitor progress. TQIntelligence is working on developing technology that uses voice samples to detect the severity of mental health conditions in children from low-income communities.
Dec 09, 2021 2,445 words in the original blog post.
Bradley Metrock, CEO of Project Voice, presented the opening keynote at Project Voice X. He discussed how the landscape for voice technology and conversational AI has changed since the pandemic, with many players gone and new interest emerging. Metrock highlighted three key points: contact centers are becoming essential in various industries; there is a growing acceptance of voice payments due to changes in purchasing habits caused by the pandemic; and there will be a shortage of skilled professionals in the field, leading to increased merger and acquisition activity. He also shared a personal story about Leland, who underwent a hemispherectomy for his condition, emphasizing the importance of making hard cuts and choices for growth and success.
Dec 08, 2021 3,145 words in the original blog post.
In this tutorial, the author guides readers through creating, publishing, installing, and using their first npm package. The process involves setting up a unique name for the package, version control, exporting functions or classes, and testing the package locally before publishing it. Additionally, the author demonstrates how to use an existing npm package in another project by requiring it and utilizing its exported functions or classes.
Dec 06, 2021 1,170 words in the original blog post.
This tutorial demonstrates how to build a searchable phone call dashboard using Twilio and Deepgram's Speech Recognition API. The project involves setting up an Express server, creating a database for storing call information, and building a front-end interface with Vue.js to enable users to search through transcriptions of recorded calls. The final code is available on GitHub at https://github.com/deepgram-devs/twilio-voice-searchable-log.
Dec 02, 2021 1,900 words in the original blog post.