August 2023 Summaries
26 posts from Deepgram
Filter
Month:
Year:
Post Summaries
Back to Blog
The International Conference on Machine Learning (ICML) is an annual event where the brightest minds in AI come together to share their latest research and technological breakthroughs. In 2023, it was held in Waikiki, with a focus on topics such as Reinforcement Learning with Human Feedback (RLHF), Differential Privacy, disinformation, fake news, and marginalized languages. The conference also featured discussions about the use of AI for fact-checking and the development of a robot dog using Deepgram's speech recognition software. Additionally, the Test of Time Award was given to a paper on learning fair representations that has had significant impact in the field over the past decade.
Aug 31, 2023
2,035 words in the original blog post.
Automatic Speech Recognition (ASR) is a technology that transforms spoken language into written text and has been advancing rapidly in recent years. It intersects with voice recognition, machine learning, and natural language processing to improve its accuracy and contextual understanding. Despite the potential of ASR being widely recognized, many businesses still underutilize it. The history of ASR began in 1952 with Audrey, a digit recognizer, and has evolved significantly over time, incorporating machine learning and natural language processing for better accuracy and contextual relevance. Today's ASR systems are incredibly accurate, affordable, and fast, with potential applications ranging from personal voice assistants to professional transcription services and telecommunication improvements. The future of ASR is expected to be more conversational, with machines understanding the nuances of human speech and sentiment.
Aug 29, 2023
1,910 words in the original blog post.
Researchers have developed a new benchmark called API-Bank for testing how well large language models (LLMs) use external tools such as APIs to accomplish tasks. The benchmark evaluates LLMs' abilities in three main areas: deciding when to call an API, finding the right tool for the job, and employing multiple APIs to complete a task. GPT-4 outperforms GPT-3.5 Turbo on most of the tests, but both models struggle with tasks requiring multiple rounds of interdependent API calls. The results highlight the potential for LLMs to become more efficient and useful by incorporating external tools, as well as areas where further improvements are needed.
Aug 28, 2023
2,334 words in the original blog post.
Generative AI models are increasingly relying on synthetic data for training purposes due to the scarcity of high-quality real-world data. However, this practice can lead to Model Autophagy Disorder (MAD), where a model collapses after being repeatedly trained on AI-generated data. A study found that ChatGPT's performance has degraded over time as it began relying more heavily on synthetic data. This raises concerns about the quality of outputs from generative models and highlights the need for preserving original data to maintain better model performance.
Aug 25, 2023
728 words in the original blog post.
In this tutorial, the author demonstrates how to create a sentiment analysis tool using Hugging Face and Deepgram APIs. The project involves setting up a Jupyter notebook with necessary packages such as transformers, numpy, pandas, matplotlib, seaborn, scipy, deepgram-sdk, python-dotenv, torch, and pytube. The process includes transcribing audio with Deepgram, setting up the sentiment analysis pipeline with Hugging Face, converting output into one composite sentiment score, turning data into a sentiment chart, and smoothing the chart for better visualization. This tool can be used to analyze any video or audio clip's sentiments over time, providing insights into customer perceptions, brand reputation, market trends, and more.
Aug 24, 2023
1,557 words in the original blog post.
The Llama-2 paper presents a collection of four generative AI models with varying parameters, built on a classic Transformer Architecture. These models employ RMSNorm for easier handling of large token numbers and weights, SwiGLU activation function to determine neuron activity, and RoPE method to emphasize the importance of word positions in sentences. Llama-2 also has a chatbot version called "Llama-2-Chat" that uses Ghost Attention (GAtt) for improved dialogue flow over multiple turns. Comparisons show Llama-2 outperforming its predecessor, Llama-1, and other open-source models in various benchmarks. However, it does not beat closed-source models like GPT-4 and PaLM-2-L. Meta did not use any user data from their platforms for training Llama-2 and has offset all carbon emissions generated during its development.
Aug 23, 2023
1,318 words in the original blog post.
Massive Multitask Language Understanding (MMLU) is a challenging NLU benchmark developed by Hendrycks et al. to measure how well an LLM understands language and can solve problems with the knowledge it encountered during training. MMLU contains 15,908 questions from various subjects at varying depths, testing qualitative and quantitative analysis, knowledge about human behavior and society, empirical methods, fluid intelligence, and procedural knowledge. The benchmark is scored by averaging each model's performance per category and then averaging these four scores for a final score. MMLU has revealed intriguing insights into LLM performance across different subjects and continues to be a valuable tool in identifying specific areas where LLMs may need improvement.
Aug 22, 2023
972 words in the original blog post.
In this article, Brad Nikkel discusses TruthfulQA, a benchmark designed by Lin et al. in 2021 to evaluate the truthfulness of large language models (LLMs) when answering questions. The benchmark consists of 817 diverse questions spanning various categories such as health, law, finance, and politics. Unlike other LLM benchmarks like ARC, HellaSwag, and MMLU, TruthfulQA focuses on measuring the truthfulness of LLM outputs rather than their ability to reason or understand language.
TruthfulQA assigns a truth score between 0 and 1 for each statement based on its probability of being true. The benchmark evaluates both the truthfulness and informativeness of LLM-generated answers, with human judges scoring each machine-generated answer. In addition to the main task, there is also a secondary task where LLMs are asked to pick answers from multiple-choice questions (some true and some false), which are then scored automatically.
The results showed that while larger models like GPT-3-175B were more informative, they were less truthful compared to smaller models. This suggests that scaling alone may not be enough to address LLMs' truth deficits, and fine-tuning and prompt engineering might be necessary for creating more truthful models.
TruthfulQA has contributed significantly to the field by highlighting the challenges of designing LLMs that generate relevant and true responses. It serves as a reminder that novel benchmarks will continue to emerge alongside advancements in LLM technology, driving improvements in language modeling.
Aug 22, 2023
1,192 words in the original blog post.
HellaSwag is a large language model (LLM) benchmark designed by Zellers et al. in 2019 to evaluate commonsense reasoning in LLMs. The dataset tests common-sense natural language inference (NLI) about physical situations and uses adversarial filtering to generate deceptive, challenging incorrect answers for a multi-choice test setting. When initially released, state-of-the-art models like BERT had poor commonsense reasoning, with human accuracy soaring above 95% while these cutting-edge models mustered accuracies below 50%. Since its release, HellaSwag has pushed the field to evolve benchmarks and improve LLM performance.
Aug 21, 2023
835 words in the original blog post.
This post delves into deep model pruning, distillation, and quantization techniques that help address the challenges posed by increasing complexity and resource demands of modern neural networks. These methods aim to reduce model size and improve efficiency, enabling deployment on a wide range of devices and opening up possibilities for real-world applications across various domains. The post covers the principles behind deep model pruning, distillation, and quantization in detail, outlines the steps of the processes, and discusses the trade-offs involved.
Aug 21, 2023
9,965 words in the original blog post.
A study by Humboldt University research assistant Jennifer Haase and University of Essex psychology lecturer Dr. Paul Hanel compared human and AI-generated ideas' creativity using the Alternate Use Task (AUT). They found that Generative Artificial Intelligence (GAI) chatbots, such as Alpa.ai, Copy.ai, ChatGPT3, Studio, and YouChat, were about as creative (or uncreative) as most humans in generating novel uses for everyday objects. The researchers argue that creativity can be defined simply as "creating something new and useful," which aligns with our common-sense notion of creativity. While some chatbots may have seen AUTs in their training data, the study found no significant difference in originality among the five chatbots tested.
Aug 18, 2023
1,870 words in the original blog post.
Every two weeks, a language dies, with over 7000 spoken worldwide, many linguists predict that at least half will be extinct within the next century. Language preservation is crucial as it carries the knowledge and history of cultures. The decline in minority languages can be attributed to globalization, migration, and technology's incentives to communicate in dominant languages. However, AI technologies are offering a way to preserve endangered languages from extinction. In Nigeria, which has over 500 spoken languages, AI is being used to build language tech like automatic speech recognition (ASR) and speech-to-text tools. While there are challenges such as lack of resources and ethical considerations, organizations like Masakhane are working towards strengthening NLP research in Africa by offering tools for training models for various African languages.
Aug 17, 2023
1,089 words in the original blog post.
The text discusses the process of AI language detection in automated speech recognition (ASR) applications. It explains how machine learning is used to solve a classification problem where the objective is to accurately identify the label or language of a given text or audio sample. Features such as Mel-frequency cepstral coefficients (MFCCs) are extracted from training data and passed to a model for training, which learns to identify underlying patterns and correlations between the features and corresponding language labels. The Deepgram API is used to perform language detection on audio samples, and the text demonstrates how this works in real-time using conlangs or constructed languages. The author also provides code examples of how to use the Deepgram Python SDK for language detection.
Aug 16, 2023
2,036 words in the original blog post.
The ARC Benchmark is a challenging test for large language models (LLMs) that focuses on their reasoning abilities and knowledge in answering questions. Developed by Clark et al. in 2018, the AI2 Reasoning Challenge benchmark aimed to push LLMs beyond simple fact retrieval tasks and evaluate their ability to answer complex, multi-faceted questions requiring reasoning, commonsense knowledge, and deep comprehension skills. The ARC dataset contains 7787 non-diagram, multiple-choice science questions split into an "Easy Set" and a "Challenge Set." The Challenge Set includes questions that stumped retrieval-based and word co-occurrence algorithms, making them more difficult for LLMs to answer. ARC also offers the ARC Corpus, which contains 14 million sentences relevant to the questions in the dataset, designed to help models solve ARC questions without outright memorizing answers. The benchmark has been used to evaluate various LLMs and continues to be a valuable tool for assessing their question-answering capabilities.
Aug 15, 2023
1,021 words in the original blog post.
Google has introduced an AI tool for journalists that aims to assist in generating news articles and improving productivity without replacing human reporters. This move is part of a growing trend among news organizations incorporating AI into their work processes, with tools like ChatGPT and Jasper AI automating repetitive tasks and streamlining the writing process. AI can also be useful for analyzing large amounts of data in data journalism. As AI becomes more prevalent in newsrooms, it is essential to learn how to use these tools responsibly and effectively without sacrificing journalistic principles. Some ethical considerations include addressing potential biases in AI-generated content and establishing accountability for errors or missteps in AI-generated work.
Aug 14, 2023
1,168 words in the original blog post.
"Calling Your Video Game With Your Phone" is a 3-part series that demonstrates how to use Twilio and Deepgram to create phone calls which patch into your video game via a websocket server. In this first part, the author explains how to use Twilio to stream audio to a server, forward it to Deepgram for transcription, and then forward the results to a game instance. The example game is written in Godot, but can be adapted to other platforms like Unity or non-game front-end applications. The server used here is written in Rust, but understanding the core architecture of the system should allow you to build analogous servers in other languages.
The series aims to provide a foundation for building more complex applications where phone calls can impact gameplay and create unique gaming experiences. In future parts, the author plans to expand on this server to enable interactions such as having actual conversations with characters in a game or ordering virtual items via phone calls.
Aug 11, 2023
2,283 words in the original blog post.
In Part 2 of the "Calling Your Video Game With Your Phone" series, the author extends the server from Part 1 to allow a game client to communicate with the player via Text-to-Speech (TTS). This enables more interactive experiences between the game and the phone, such as having conversations with in-game NPCs. The TTS service used here is Amazon Polly, but the strategy can be adapted for any other TTS provider. The server architecture has been modified to accommodate this new functionality, allowing the game client to send text messages back to the server, which then converts them into audio and sends them to Twilio for playback on the caller's phone. A simple Godot game demonstrates how this can be implemented in a game engine. The author also provides an overview of the code changes made to the server from Part 1 to enable TTS functionality.
Aug 11, 2023
1,951 words in the original blog post.
In the final part of the 3-part series, we explore Robot Dreams, a game made for Godot Wildjam #55. The game uses a modified server from Part 2 to offer a unique gameplay twist where players can call their video game with their phone. Technical changes in the server include the ability to specify a config file via the command line and the use of an AWS instance with docker and docker-compose for deployment. While Robot Dreams has some shortcomings, such as lack of feedback on the phone and limited language support, it demonstrates how integrating voice technology can expand accessibility in games.
Aug 11, 2023
808 words in the original blog post.
SuperGLUE is a more complex benchmark for evaluating Language Models (LLMs) compared to the GLUE benchmark introduced in 2019. It offers a new set of tasks, as well as a public leaderboard for assessing language models' performance. The SuperGLUE benchmark includes eight subtasks and two additional "metrics" that analyze the model at a broader scale. These tasks are designed to be solvable by an English-speaking college student but surpass what current (late 2019) language models can accomplish. The final SuperGLUE benchmark score is computed as the simple average across all tasks. Unlike HuggingFace leaderboard for LLMs, the leaderboard for SuperGLUE is populated mainly by models developed by smaller research labs rather than well-known close-sourced models such as Claude and GPT.
Aug 09, 2023
1,208 words in the original blog post.
The article discusses the importance of language model benchmarks (LLMs) in evaluating AI performance, particularly large language models like GPT-4. These benchmarks provide an objective measure for developers and users to compare competing models based on their ability to complete specific natural language processing tasks. They also offer valuable insights into areas where a model excels or struggles, helping researchers gauge the current state of the art in AI research.
The history of AI and LLM benchmarks is traced back to early machine translation systems in the 1960s-70s, followed by bag-of-words models in the 1980s-90s, sequence models and named entity recognition in the early 2000s, word embeddings in the mid-2010s, attention models and question answering in the late 2010s, and finally, GLUE and SuperGLUE benchmarks. The article also highlights some emerging trends in LLM benchmarking, such as a focus on ethical aspects like fairness and bias, explainability, and expanding capabilities beyond basic NLP tasks.
The author emphasizes that no single test can fully capture an LLM's wide array of abilities and potential weaknesses, making comprehensive benchmarking crucial for understanding these complex AI systems. The article concludes by providing a list of Deepgram articles covering various LLM benchmarks, with plans to update the list as new benchmarks emerge.
Aug 09, 2023
2,556 words in the original blog post.
Large Language Models (LLMs) are being used to create autonomous "agents" capable of handling complex tasks with minimal human input. These LLM agents can manipulate their environment using tools such as APIs, and they have the potential to revolutionize various industries by automating tasks that previously required human intervention. Several experimental LLM agent-based projects are gaining momentum, with some predicting that Artificial General Intelligence (AGI) will likely involve some flavor of agent framework. Academic research on LLM agents includes projects like Voyager, which learned how to interact with Minecraft's 3D world and continuously improved its skills through exploration and experimentation. Independent open-source LLM agent projects include AutoGPT, BabyAGI, GPT Researcher, and others that leverage existing LLM frameworks or create custom agents from scratch. While there are challenges to overcome in the development of autonomous goal-driven LLM agents, such as reliable translation between natural language and actions, decreasing inference costs, and increasing tool availability, these issues do not seem insurmountable. The potential for LLM agents to automate complex tasks and make computing more accessible to a wider range of people is an exciting prospect that could transform various industries.
Aug 08, 2023
3,675 words in the original blog post.
The article discusses synthetic data, which is generated using machine learning algorithms and helps bypass privacy laws while still providing useful training datasets for AI models. Synthetic data can be created for any type of dataset, from simple tabular data to complex unstructured data, using various techniques such as Variational Auto-Encoders (VAE), Generative Adversarial Networks (GAN), and Diffusion Models. The use of synthetic data is particularly valuable in industries with privacy concerns or limited access to quality data, such as healthcare, finance, and AI research. Synthetic data can help build more realistic language models by providing high-quality training data and addressing biases present in real-world datasets. However, the reliance on real-world data for generating synthetic data raises concerns about maintaining privacy and ensuring accurate representation of the original data.
Aug 07, 2023
1,159 words in the original blog post.
Machine Unlearning is a concept that aims to make a model forget or unlearn specific portions of its training dataset, serving as the converse of machine learning. This technique has gained importance due to the "Right to be Forgotten" legislation under the European Union's General Data Protection Regulation (GDPR). While it holds promise in complying with regulations and rectifying factually incorrect information within models, challenges such as determining the effectiveness of unlearning, ensuring complete data point forgetting, and quantifying the exact influence of data points on the model persist. Existing algorithms for machine unlearning are classified into two main categories: exact and approximate unlearning methods. The outlook for machine unlearning is promising, with potential applications in maintaining privacy rights while respecting individual artists' works. However, addressing challenges will be a critical part of developing effective and efficient machine unlearning algorithms.
Aug 04, 2023
962 words in the original blog post.
The article discusses eight major conversational AI conferences scheduled for 2023 and beyond. These conferences are designed to bring together professionals in the field of Conversational AI, allowing them to share knowledge, network, and learn about new developments in the industry. Some notable conferences mentioned include Project Voice, ACM Conversational User Interfaces (CUI), Ai4, Voice & AI 2023, AI Hardware & Edge AI Summit, The AI Conference SF, World AI Summit, ExCel London Chatbot Summit, and Ibero America Chatbot & Conversational AI Summit. These conferences are expected to attract thousands of attendees from around the world, including representatives from major tech companies like Google, Microsoft, Apple, and Amazon.
Aug 03, 2023
1,352 words in the original blog post.
Deepgram introduces a new feature called Filler Words that transcribes filler words and disfluencies found in English audio, both pre-recorded and streaming. This feature is compatible with existing features like Smart Formatting and Diarization and initially supports the Nova general speech-to-text model. The Filler Words feature has no impact on latency or performance and consistently spells disfluencies throughout the transcript. It caters to customers who require precise, verbatim transcripts for various use cases such as sales enablement, public speaking coaching, legal transcription, and more. To use this new feature, users need to set tier=nova&model=general&filler_words=true when using the English Nova general model. Deepgram encourages feedback on their features and services through GitHub discussions or contacting their product experts for further information.
Aug 02, 2023
406 words in the original blog post.
The article discusses why bigger isn't always better for language models in AI. It highlights how OpenAI's GPT-4 model, with over 1.7 trillion parameters, is not necessarily superior to smaller alternatives like Falcon 40B-instruct and Alpaca 13B. The article argues that larger models are more expensive to train and deploy, harder to control and fine-tune, and can exhibit counterintuitive performance characteristics. It also points out that users often seek alternatives that are less costly and better suited for their needs. Furthermore, the article mentions how smaller language models can be trained using imitation learning techniques from larger models like GPT-4, offering a more balanced mix of performance, cost, and usability.
Aug 01, 2023
1,807 words in the original blog post.