Home / Companies / Gladia / Blog / Post Details
Content Deep Dive

Automatic Speech Recognition (ASR): how speech-to-text models work and which one to use

Blog post from Gladia

Post Details
Company
Date Published
Author
Anna Jelezovskaia
Word Count
3,422
Company Posts That Month
19
Language
English
Hacker News Points
-
Post removed?
No
Summary

Automatic speech recognition, or speech-to-text, converts spoken audio into written language by combining an audio encoder that interprets sound with a language model that generates coherent text. Modern systems are generally grouped into encoder-decoder, CTC, encoder-transducer, continuous-input speech LLM, and discrete-input speech LLM architectures, each differing in how audio and language components interact and in their trade-offs between accuracy, latency, streaming capability, scalability, and flexibility. Models such as Wav2Vec2, Whisper, Kyutai-STT, and NVIDIA’s Nemotron-Speech-Streaming illustrate these approaches, from self-supervised CTC systems and robust large-scale encoder-decoders to low-latency streaming designs. The discussion emphasizes that no model is universally best: organizations should evaluate candidates using their own audio, languages, accents, noise conditions, real-time requirements, deployment preferences, and total operating costs rather than relying solely on public benchmark scores such as word error rate.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 37 747 162 79 -85%
Vector Search 19 265 57 33 -89%
Real-time 13 649 155 80 -85%
Voice AI 2 324 41 16 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.