Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Francesco, Ivan Medennikov, Taejin Park, Tatiana Timofeeva, Jagadeesh Balam, Adi- margolin, and Maryam Motamedi
Word Count
3,324
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

NVIDIA’s Nemotron 3 Diarization is an open-weight, 100-million-parameter speaker diarization model designed to identify when up to eight anonymous speakers are active in live or recorded 16 kHz mono audio, including overlapping speech. Built on the Sortformer approach, it maintains consistent arrival-ordered speaker channels across streamed audio chunks using speaker-cache and recent-context mechanisms, with configurable input-buffer latency from 30.4 seconds to a recommended minimum of 0.32 seconds. NVIDIA reports that the model ranked first in VoiceArena’s initial Diarization-Bench, achieving a 14.72% diarization error rate across 139 English conversations, and showed lower error rates and substantially higher batched throughput than its prior four-speaker Streaming Sortformer baseline across several datasets. The model produces speaker activity timestamps rather than identities or transcripts, so applications can combine its output with automatic speech recognition and external metadata to create speaker-attributed transcripts. The article also provides NeMo-based setup and inference guidance, discusses latency, accuracy, speaker-count, and acoustic limitations, and advises developers to test complete pipelines on representative audio, especially for consequential uses.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 22 649 155 80 -85%
Vector Search 1 265 57 33 -89%
Voice AI 1 324 41 16 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.