Home / Companies / Speechmatics / Blog / Post Details
Content Deep Dive

Best-in-class real-time ASR system

Blog post from Speechmatics

Post Details
Company
Date Published
Author
Steve Kingsley
Word Count
1,085
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Real-Time ASR Systems have two common modes of operation: batch and real-time. In batch mode, audio is provided in complete files with a single transcript output, allowing higher accuracy. Real-time systems provide an audio stream, returning short segments of transcription back at regular intervals, where the trade-off between latency and accuracy comes into play. Evaluating batch versus real-time ASR, Ursa outperforms competitors in accuracy even when prioritizing speed over accuracy, achieving near-batch levels of accuracy with low latency settings. Latency is controlled through `max_delay` and `max_delay_mode`, allowing for a balance between timeliness and accuracy. Ursa's latest release demonstrates outstanding performance, reducing to zero relative difference in WER as latency increases to 10s, outperforming major vendors such as Amazon, Microsoft, and Google.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 20 2,062 598 178 +12%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.