Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Word error rate is broken: How to actually evaluate speech-to-text in 2026

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
3,470
Company Posts That Month
44
Language
English
Hacker News Points
-
Post removed?
No
Summary

In 2026, the traditional Word Error Rate (WER) metric for evaluating speech-to-text systems is deemed inadequate due to its inability to account for contextual accuracy and its tendency to penalize more accurate models. AssemblyAI's workshop highlighted how WER, which treats all word errors equally, often misrepresents a model's true performance, especially when models like Universal-3 Pro correctly transcribe words that human transcribers miss. The workshop demonstrated that WER's limitations are particularly problematic in sectors where specific terminology is crucial, such as medical and legal fields. New evaluation metrics like Semantic WER, Missed Entity Rate (MER), and LLM-as-a-Judge (LASER) scoring offer more nuanced assessments by considering domain-specific word lists, the importance of named entities, and semantic meaning preservation. These metrics provide a more comprehensive evaluation framework, reflecting real-world transcription accuracy and guiding industry practices towards a multi-metric approach tailored to specific use cases, thereby enhancing the reliability and effectiveness of speech-to-text technology.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 14 5,932 1,046 223 -2%
Real-time 12 6,296 1,346 246 -2%
Voice AI 8 2,379 221 38 -3%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.