November 2023 Summaries
16 posts from AssemblyAI
Filter
Month:
Year:
Post Summaries
Back to Blog
In this tutorial, we learn how to use LLMs (Large Language Models) to extract insights from a customer call in just a few lines of code. We first transcribe the audio file using AssemblyAI's Python SDK and then utilize LeMUR, their framework for building LLM applications on audio data, to extract a summary, action items, and contact information. The prompts provided ensure that the LLM accurately identifies key details from the call transcript. This method can be applied across various industries to gain valuable insights from customer interactions and improve services accordingly.
Nov 30, 2023
938 words in the original blog post.
Customer conversational data, often underutilized, can be analyzed using AI-powered call analytics tools that leverage Speech AI models such as Automatic Speech Recognition, Audio Intelligence, and Large Language Models. These advanced technologies help companies automate previously manual processes for analyzing customer interactions at scale, leading to more efficient quality monitoring, identification of trends, and increased agent performance and productivity. By integrating AI-powered call analytics, businesses can unlock insights from conversational data that inform strategic decisions about training, branding, and customer satisfaction.
Nov 30, 2023
885 words in the original blog post.
Speech AI models are transforming the field of telehealth by improving communication between patients and providers, automating tasks, and providing additional insights for personalized care. These AI models can transcribe conversations in real-time, enhance online therapy sessions, summarize appointments, analyze patient experience, review patient's history, and even act as a virtual nursing assistant. HIPAA compliance and data privacy measures are crucial considerations when integrating Speech AI into telehealth platforms. The benefits of incorporating Speech AI include time savings for administrative tasks, prevention of employee burnout, and improved patient experience with accurate diagnosis through well-documented patient history.
Nov 30, 2023
1,225 words in the original blog post.
The text discusses AI speech recognition systems, their benefits, use cases, and the top considerations when deciding whether to build or buy such a system. It highlights that while companies may be tempted to build their own AI speech recognition models, pre-existing models from an AI partner often have better accuracy due to large diverse datasets and ongoing updates based on AI research. Additionally, building an in-house model requires extensive internal resources and expertise, as well as addressing data security concerns and compliance requirements. The text also provides a checklist for companies considering the build or buy decision for an AI speech recognition system.
Nov 27, 2023
1,397 words in the original blog post.
This week, AssemblyAI introduces custom text input for LeMUR, allowing users to submit their own custom text inputs without needing a transcript ID. They also announce their participation at AWS re:Invent and partnership with Amazon Web Services (AWS) Marketplace. A revamped Playground is launched, offering improved user experience and combined features of Async and LeMUR Playgrounds. The blog section features tutorials on getting YouTube video transcripts, real-time transcription in Python, obtaining Zoom transcripts using the Zoom API, and more.
Nov 24, 2023
375 words in the original blog post.
AssemblyAI has become a partner on the Amazon Web Services (AWS) Marketplace, providing users with easy access to powerful Speech AI applications for various use cases. The partnership allows companies to offset AWS committed spend by investing in AssemblyAI's Speech AI models. AWS will contribute 100% of marketplace purchases to unused EDP commitment, and all products sold on the marketplace must pass Amazon's due diligence and security validations. AssemblyAI is also participating in the AWS Generative AI Pavilion at AWS re:Invent from November 27 - December 1, 2023, showcasing its innovative technology and offering networking opportunities for attendees.
Nov 20, 2023
308 words in the original blog post.
AssemblyAI recently celebrated hitting 100K subscribers on YouTube, highlighting a variety of popular videos since starting the channel. They've also introduced LeMUR, a framework for applying Language Models to audio data and provided tutorials on how to build LangChain Audio Apps with Python in five minutes and integrate audio into LangChain.js apps in five minutes. AssemblyAI is participating at AWS re:Invent from November 27-30th where they will discuss their collaboration with Google Cloud TPUs, which enhances speech recognition capabilities. This collaboration also provides better cost efficiency, advances the Conformer-2 model for large-scale testing and ensures smooth compatibility across platforms. Furthermore, AssemblyAI has released new blog posts on real-time transcription in Python, automatically determining video sections with AI using Python, and how to get Zoom Transcripts using the Zoom API using Python.
Nov 16, 2023
440 words in the original blog post.
To successfully integrate AI into your business, follow these steps:
1. Identify user value.
2. Define clear and measurable goals.
3. Choose a model that aligns with your goals.
4. Consider whether to build in-house or collaborate with a third-party AI partner.
5. Identify the KPIs.
6. Create cross-functional teams.
7. Ship, maintain, and iterate.
Potential blockers and risks include difficulty in monetizing the product, limited developer and engineer resources, and keeping up with the rapidly changing field of AI.
By following these steps and best practices, businesses can successfully leverage AI to improve their products and better serve their customers.
Nov 15, 2023
2,115 words in the original blog post.
Text Summarization is a technique used to shorten the length of text documents while maintaining their most important information. It can be applied to different types of content such as articles, emails, and meeting transcripts. There are two main approaches to Text Summarization: extractive and abstractive. Extractive methods select and copy parts of the original text into the summary, while abstractive methods generate a new, condensed version of the text that conveys its most important points. Both techniques have their own advantages and limitations, with abstractive summarization generally considered to be more advanced but also more challenging to achieve high quality results. In recent years, many Text Summarization APIs and AI models have been developed by various companies and organizations, including AssemblyAI, plania, Microsoft Azure, and MeaningCloud. These tools can be used in a wide range of industries and applications, such as meeting transcript summarization, video chapter generation, and podcast editing.
Nov 09, 2023
2,447 words in the original blog post.
New models for Punctuation Restoration and Truecasing have been introduced, outperforming previous production models on various data and metrics. The new models show significant improvements in handling casing for challenging linguistic types such as mixed-case words (+39% F1 score), acronyms (+20% F1 score), and capital-case (+11% F1 score). Overall, there is a 17% relative improvement on average across test datasets for predicting upper-case letter classification. Punctuation accuracy improves by 11% (F1 score). The new models are already in production, with API users automatically benefiting from the upgrades.
Nov 08, 2023
1,759 words in the original blog post.
In this tutorial, we learned how to use AssemblyAI's Auto Chapters model in Python to automatically generate chapter titles and timestamps for a YouTube video from its transcript. We also demonstrated how to improve the names of the video sections using LeMUR, an open-source large language model. This method can help content creators enhance their viewers' experience by making it easier for them to navigate and find specific information within a video.
Nov 07, 2023
1,579 words in the original blog post.
In this week's update, we introduce new Punctuation Restoration and Truecasing models that have shown significant improvements in accuracy across various metrics. These enhancements are expected to benefit users by providing more accurate casing and punctuation handling in text data. We also announce our participation at AWS re:Invent from November 27th to 30th, where we'll be available for discussions and demonstrations of our AI solutions. Additionally, our Speech-to-Text documentation has been revamped with new features such as Summarization, Sentiment Analysis, Auto Chapters, and PII Redaction. Lastly, our YouTube channel is nearing 100k subscribers, and we share some popular tutorials on topics like machine learning, OpenAI's API in Python, vector databases, key phrase detection, video sectioning, and AI-based audio summarization.
Nov 07, 2023
462 words in the original blog post.
AssemblyAI offers state-of-the-art Text Summarization models that are fast, scalable, and continuously updated by an in-house team of AI experts to keep it state-of-the-art as new research emerges. These models are accessible through a single API call, making it easy for teams of all sizes to embed the models into their products. The Text Summarization models can be used across various industries and use cases, including Conversation Intelligence, Video/Media Platforms, Podcasts, and Virtual Meeting Platforms.
Nov 03, 2023
1,397 words in the original blog post.
This article discusses the functioning of Speech-to-Text AI, also known as Automatic Speech Recognition (ASR), and key considerations when selecting a suitable technology. Modern speech-to-text methods mainly involve End-to-End Deep Learning to route an acoustic waveform into a sequence of words. The accuracy of the transcriptions depends heavily on the training provided to the AI model with large amounts of data. Key aspects to consider include near human-level accuracy, features like automatic punctuation and speaker diarization, noise robustness, confidence scores, language support, consistent innovation, scalability, and security measures. The choice between free or paid plans is also crucial in determining the best fit for a business's needs.
Nov 03, 2023
1,093 words in the original blog post.
In this tutorial, you learned how to use the AssemblyAI Python SDK to transcribe an audio file and detect key phrases within it. Here's a step-by-step breakdown of what you did:
1. Set up your virtual environment for Python and install the necessary dependencies.
2. Get your AssemblyAI API token from the dashboard on their website, and save it in an environment variable.
3. Import the required modules and classes from the AssemblyAI SDK.
4. Define a function to upload an audio file to the AssemblyAI platform using the transcription endpoint of their API.
5. Use this function to transcribe your audio file and print out the JSON-formatted response from the server.
6. Create another function to download the transcripted text in plain format from the server.
7. Call this function, passing in the filename of your audio file, and store the resulting transcript in a variable.
8. Define a third function to detect key phrases within the transcribed text using the auto_highlights attribute of the SpeechRecognitionResult class.
9. Use this function to extract the highlights from your transcript, sort them by timestamps if desired, and print out their content along with relevant metadata like rank and count.
By following these steps, you can easily analyze audio data for key phrases using Python and the AssemblyAI platform.
Nov 02, 2023
1,037 words in the original blog post.
This week, AssemblyAI introduced a number of product improvements, including faster transcription speed, clearer error messages, and an update for its Rivet integration. The company also highlighted LeMUR, a tool that enables developers to build applications using large language models (LLMs) on voice data. A variety of new cookbooks were shared to demonstrate the capabilities of LeMUR, such as sentiment analysis in customer calls and generating action items from meetings. Additionally, AssemblyAI's blog featured articles on integrating speech recognition and diarization into a unified model for multi-speaker processing, using audio data with Python in LangChain, and converting speech to text in Python in five minutes. Finally, the company shared tutorials on creating AI agent teams with AutoGen, building an AI Audio Bot, and utilizing retrieval augmented generation on audio data with LangChain.
Nov 01, 2023
374 words in the original blog post.