Diarization error rate (DER) explained
Blog post from Gladia
Diarization Error Rate (DER) measures how accurately a system identifies who spoke when by combining missed speech, false alarms, and speaker-confusion time relative to ground-truth speech duration, and it can exceed 100% because these components are additive. While Word Error Rate evaluates transcription accuracy, DER is essential for multi-speaker recordings because correct words assigned to the wrong person can undermine meeting summaries, CRM records, sentiment analysis, and automated quality assurance. The discussion also distinguishes word-level and token-level diarization error rates, which may better reflect the impact of attribution errors on LLM-based workflows, and recommends monitoring Jaccard Error Rate for conversations with unequal speaker participation. Performance varies substantially with real-world conditions including overlapping speech, noise, reverberation, speaker similarity and count, mono versus stereo capture, and narrowband telephony audio; stereo channels and accurate speaker-count constraints can reduce confusion. It recommends evaluating systems with representative annotated production audio using RTTM files, UEM evaluation regions, and a roughly 250-millisecond collar around turn boundaries, with DER below 15% presented as a practical target for reliable speaker-labeled analytics and below 10% for controlled, high-quality audio.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.