Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Does Whisper do speaker diarization? Whisper + pyannote, and its limits

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
2,546
Company Posts That Month
36
Language
English
Hacker News Points
-
Post removed?
No
Summary

Whisper is an automatic speech recognition model that transcribes audio but does not identify speakers, so users commonly combine it with pyannote.audio or WhisperX through a three-stage process of transcription, diarization, and timestamp-based word-to-speaker alignment. While this open-source approach can suit offline, research, or hobby projects, it requires gated-model credentials, GPU infrastructure, voice-activity detection tuning, and custom alignment logic, and can struggle with short interjections, overlapping speech, and speaker merging. The piece argues that concatenated minimum-permutation word error rate (cpWER), which measures transcription and attribution errors per speaker, better reflects practical diarization quality than diarization error rate (DER), which may understate errors in brief or overlapping turns. It presents AssemblyAI’s Universal-3.5 Pro as a managed alternative that jointly generates transcripts and speaker labels through an API, claiming lower average cpWER than several competitors, simpler setup, and per-hour pricing, while acknowledging that it is not appropriate for offline or air-gapped use.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Voice AI 3 2,839 275 56 -36%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.