Home / Companies / Deepgram / Blog / Post Details
Content Deep Dive

The Rise of PER: A New Standard for Measuring TTS Accuracy

Blog post from Deepgram

Post Details
Company
Date Published
Author
Jose Nicholas Francisco
Word Count
2,023
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Pronunciation Error Rate (PER) is presented as a phoneme-level metric for evaluating text-to-speech accuracy in areas where broad measures such as Mean Opinion Score (MOS) and Word Error Rate (WER) can conceal costly errors involving medical terms, brand names, financial language, and alphanumeric identifiers. PER calculates phoneme substitutions, deletions, and insertions against a reference pronunciation, enabling teams to identify specific sound-level failures, while forced alignment is described as more reliable for precise evaluation than ASR round-trip testing, which may generate substantial false alarms. Production implementations combine custom pronunciation dictionaries, SSML phoneme controls, automated WER and Keyword Error Rate monitoring, periodic MOS testing, and business-specific thresholds, such as under 5% KER for critical healthcare terminology or above 90% accuracy for contact-center vocabulary. The approach also emphasizes testing providers with domain-specific corpora, monitoring pronunciation alongside latency, setting alerts for degradation, and using captured failures to refine dictionaries and validate fixes before deployment.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Voice AI 4 324 41 16 -89%
Vector Search 1 265 57 33 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.