Home / Companies / Speechmatics / Blog / Post Details
Content Deep Dive

How to Effectively Introduce External Language Models in the RNNT by Subtracting Internal Language Model Scores

Blog post from Speechmatics

Post Details
Company
Date Published
Author
Oliver Parish
Word Count
1,338
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

Integrating external language models (LMs) into automatic speech recognition (ASR) systems can significantly improve accuracy while maintaining fast runtimes and low memory costs. Recent end-to-end models, such as the recurrent neural network transducer (RNNT), include an implicit internal language model (ILM) that is trained jointly with the rest of the network. However, when combining external LMs with RNNT models, it's essential to subtract the ILM scores to get the best accuracy results. This involves applying Bayes' rule and using scale factors λ₁ and λ₂ to balance the trade-off between acoustics and common word sequences. Approximating the ILM can be done through various methods, including removing acoustic data contributions or modeling it with an LSTM. Tuning the parameters λ₁ and λ₂ requires careful consideration of the dataset used, as they can significantly impact performance. By carefully balancing the contributions of external LMs and ILMs, ASR systems can achieve significant improvements in accuracy while maintaining fast runtimes.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 8 111 32 19 -15%
Real-time 3 1,351 420 143 -5%
Vector Search 1 336 71 38 +26%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.