Whisper-v3 Hallucinations on Real World Data
Blog post from Deepgram
Whisper-v3, the latest version of OpenAI's automatic speech recognition (ASR) model, has been found to hallucinate more frequently than its predecessor, Whisper-v2, when tested on real-world data. The median Word Error Rate (WER) for Whisper-v3 is 53.4, while Whisper-v2 only has a median WER of 12.7. Users have reported hallucinations in languages like Japanese and Korean as well. The author of this text tested the model on various audio files and found that it performs well with edge cases but struggles with real-world data, leading to high error rates.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 1 | 212 | 56 | 20 | +75% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.