Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Universal-2 vs OpenAI's Whisper: Comparing Speech-to-Text models in real-world use cases

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Patrick Loeber
Word Count
2,446
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

This article compares the performance of four Speech-to-Text models - Universal-2, Universal-1, Whisper large-v3, and Whisper turbo - in real-world scenarios. The evaluation focuses on proper nouns, alphanumerics, text formatting, and hallucinations. Universal-2 outperforms the other models in most categories, showing significant improvements over its predecessor, Universal-1. It has the best overall accuracy (6.68% WER), superior proper noun handling (13.87% PNER), and best formatting accuracy (10.04% U-WER). Whisper large-v3 shows some notable strengths and limitations, with the best alphanumeric transcription accuracy (3.84% WER) but also a documented propensity for hallucinations. The article concludes that Universal-2 is the leading model in most categories, offering significant improvements over its predecessor and showing a 30% reduction in hallucination rates compared to Whisper large-v3.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 2,876 370 130 -20%
TPUs 1 8 4 3 +100%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.