Entity accuracy in speech-to-text: why word accuracy isn't enough
Blog post from AssemblyAI
Speech-to-text systems can achieve low word error rates while still failing on the names, phone numbers, account codes, emails, locations, and medical terminology that determine whether real-world workflows succeed, so the piece argues that Missed Entity Rate (MER) is a more useful production metric. It recommends defining entity categories relevant to a use case, scoring them separately against reference transcripts, and testing models on realistic noisy, accented, or conversational audio rather than curated recordings. Citing the Pipecat benchmark, it highlights that models with relatively similar overall WER can differ substantially in entity error rates, with names described as particularly difficult. The piece also describes methods intended to improve performance, including keyterm lists, broader conversational context, agent-question context, and custom vocabulary, while warning that excessive prompting bias can reduce general accuracy. Medical transcription is presented as a high-stakes case where errors in drug names or dosages can create safety and liability risks; the article promotes AssemblyAI’s Medical Mode and related privacy features while acknowledging persistent limitations involving ambiguous speech, self-corrections, and very short utterances.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.