Introducing NVIDIA Nemotron 3.5 ASR Streaming
Blog post from Baseten
NVIDIA Nemotron 3.5 ASR is a speech recognition model available in the Baseten Model Library, designed for low-latency, high-accuracy transcription in both English and multilingual contexts, supporting 40 language locales. Featuring a cache-aware FastConformer-RNNT architecture with a 24-layer encoder and an RNNT decoder, it offers efficient streaming inference with NVIDIA Inference Microservices (NIM), achieving high concurrency and throughput on a single H100. The model maintains consistent finalization latency and time-to-first-token across increased real-time streams, crucial for interactive applications. It demonstrates competitive word error rates (WER) on benchmarks, such as achieving 2.32% WER on LibriSpeech Clean for English and an average of 8.84% WER across multiple languages on the FLEURS benchmark. Developers can further customize Nemotron ASR for specialized domains or specific accents using NVIDIA NeMo through Baseten Training, and the models are available under the OpenMDW-1.1 license.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 13 | 1,106 | 270 | 109 | -81% |
| AI Model Fine-tuning | 2 | 103 | 37 | 26 | -89% |
| AI Agents | 1 | 1,180 | 266 | 113 | -80% |
| Voice AI | 1 | 1,179 | 83 | 25 | -73% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.