Open source, multilingual transcription from NVIDIA: Nemotron 3.5 ASR
Blog post from Daily
NVIDIA has introduced Nemotron 3.5 ASR, a new multilingual speech-to-text model optimized for voice agents, alongside an English-only version, Nemotron 3 ASR. These models, notable for their low latency and open-source nature, can be hosted on personal infrastructure, offering enterprises the advantage of keeping voice data secure within their own systems. The models demonstrate impressive speed and accuracy in benchmarks, setting a new standard on the latency-accuracy Pareto frontier. They allow for configurable latency via audio processing chunk sizes, catering to both voice agent and batch transcription needs. Hosting the model independently can significantly reduce costs, with potential expenses as low as $0.05 per hour for enterprise deployment compared to typical API costs. The models can be fine-tuned for specific languages, and NVIDIA provides comprehensive tools and resources to facilitate customization, enhancing the model's utility across diverse applications.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 7 | 3,155 | 274 | 58 | -9% |
| AI Model Fine-tuning | 3 | 739 | 196 | 71 | +20% |
| LLM | 1 | 6,237 | 1,165 | 246 | -31% |
| Real-time | 1 | 5,758 | 1,361 | 266 | +0% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.