Home / Companies / Modal / Blog / Post Details
Content Deep Dive

Transcribe speech 100x faster and 100x cheaper with open models

Blog post from Modal

Post Details
Company
Date Published
Author
-
Word Count
2,301
Company Posts That Month
8
Language
English
Hacker News Points
4
Post removed?
No
Summary

Open-weight automatic speech recognition models such as NVIDIA’s Parakeet and Canary and Kyutai’s STT have recently approached proprietary services in accuracy while offering high inference speeds and features including multilingual support, timestamps, and voice activity detection. Modal evaluated NVIDIA’s English-focused Parakeet and multilingual Canary models against a proprietary transcription API using a week of ESB benchmark audio, reporting comparable or slightly better error rates alongside configurations that were either more than 100 times faster or up to 200 times cheaper, with one optimized case transcribing a week of audio in about a minute for roughly one dollar. The comparison emphasizes end-to-end throughput, including cold starts, network transfer, and data movement rather than model execution speed alone, and focuses on large-scale batch workloads rather than low-latency streaming use cases. Key engineering approaches included distributing shuffled audio across workers for balanced workloads, sorting recordings by duration within GPU batches to reduce idle processing time, using sufficiently large GPU inference batches, parallelizing file downloads, and empirically selecting GPU types and worker counts based on cost-throughput trade-offs.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 6 5,432 1,252 271 +11%
AI Model Fine-tuning 1 867 189 73 +71%
LLM 1 4,922 763 224 +11%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.