Transcribe speech 100x faster and 100x cheaper with open models
Blog post from Modal
Open-weight automatic speech recognition models such as NVIDIA’s Parakeet and Canary and Kyutai’s STT have recently approached proprietary services in accuracy while offering high inference speeds and features including multilingual support, timestamps, and voice activity detection. Modal evaluated NVIDIA’s English-focused Parakeet and multilingual Canary models against a proprietary transcription API using a week of ESB benchmark audio, reporting comparable or slightly better error rates alongside configurations that were either more than 100 times faster or up to 200 times cheaper, with one optimized case transcribing a week of audio in about a minute for roughly one dollar. The comparison emphasizes end-to-end throughput, including cold starts, network transfer, and data movement rather than model execution speed alone, and focuses on large-scale batch workloads rather than low-latency streaming use cases. Key engineering approaches included distributing shuffled audio across workers for balanced workloads, sorting recordings by duration within GPU batches to reduce idle processing time, using sufficiently large GPU inference batches, parallelizing file downloads, and empirically selecting GPU types and worker counts based on cost-throughput trade-offs.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 6 | 5,432 | 1,252 | 271 | +11% |
| AI Model Fine-tuning | 1 | 867 | 189 | 73 | +71% |
| LLM | 1 | 4,922 | 763 | 224 | +11% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.