Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Comparing End-To-End Speech Recognition Architectures in 2021

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Michael Nguyen
Word Count
3,372
Company Posts That Month
1
Language
English
Hacker News Points
6
Post removed?
No
Summary

AssemblyAI's research and development efforts focus on improving the accuracy of their Speech-to-Text API. They are exploring new architectures for end-to-end speech recognition, such as Listen Attend and Spell (LAS) and Recurrent Neural Network Transducers (RNNT). These models have shown production level accuracy matching or surpassing that of conventional hybrid DNN-HMM systems. The LAS model is perceived to have better accuracy than RNNT, but RNNT models are seen as having more desirable features for production use. Combining LAS and RNNT can achieve better accuracy and feature parity when compared to hybrid RNN-HMM models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 8 18 6 4 +20%
Real-time 8 825 221 87 -7%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.