Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC
Blog post from Hugging Face
IBM has released Granite Speech 5.0 TurboCTC, two compact 470-million-parameter English speech-recognition models designed for high-speed, accurate transcription, reportedly exceeding 12,600 real-time-factor throughput on an NVIDIA H200 GPU and transcribing more than 3.5 hours of audio per second with batched inference. The Apache 2.0-licensed TurboCTC model and the more extensively trained, CC-BY-NC-SA-4.0-licensed TurboCTC-NC model achieved aggregate word error rates of 5.00% and 4.85%, respectively, on public OpenASR tests, while also ranking among the fastest models on the far-field FFASR leaderboard. Unlike earlier Granite Speech systems that combined acoustic encoders with language models, the new encoder-only models use 16 Conformer blocks, CTC training, chunkwise attention, self-conditioning, and aggressive temporal subsampling to reduce output generation to 12.5 tokens per second, improving speed and memory efficiency but omitting features such as speech translation and keyword biasing. Both models were trained on a mix of public natural-speech datasets and synthetic multi-speaker and formatting-focused audio, with the noncommercial version additionally using GigaSpeech and SPGI Speech, and they are supported through Hugging Face Transformers for speech-to-text deployment, including on edge devices.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.