Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Entity accuracy in speech-to-text: why word accuracy isn't enough

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
2,465
Company Posts That Month
21
Language
English
Hacker News Points
-
Post removed?
No
Summary

Speech-to-text systems can achieve low word error rates while still failing on the names, phone numbers, account codes, emails, locations, and medical terminology that determine whether real-world workflows succeed, so the piece argues that Missed Entity Rate (MER) is a more useful production metric. It recommends defining entity categories relevant to a use case, scoring them separately against reference transcripts, and testing models on realistic noisy, accented, or conversational audio rather than curated recordings. Citing the Pipecat benchmark, it highlights that models with relatively similar overall WER can differ substantially in entity error rates, with names described as particularly difficult. The piece also describes methods intended to improve performance, including keyterm lists, broader conversational context, agent-question context, and custom vocabulary, while warning that excessive prompting bias can reduce general accuracy. Medical transcription is presented as a high-stakes case where errors in drug names or dosages can create safety and liability risks; the article promotes AssemblyAI’s Medical Mode and related privacy features while acknowledging persistent limitations involving ambiguous speech, self-corrections, and very short utterances.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.