How Baseten makes pyannote’s diarization models 9.6x faster
Blog post from Baseten
Pyannote’s Community-1 and Precision-2 speaker diarization models, which identify who spoke when in audio, have been optimized to improve production performance while retaining strong accuracy for applications such as meeting transcription, call analytics, and media archives. The optimizations include improved GPU scheduling and batching, reduced CPU-GPU memory transfers, fused inference passes, GPU-specific tuning, index-based clustering that reduces the cost of processing long recordings, and mixed, quality-aware numerical precision across pipeline stages. Baseten reports up to a 9.6x reduction in long-audio latency for Community-1, enabling 20 hours of audio to be diarized in about two minutes on one GPU, and about 3.2x higher Precision-2 throughput with a 0.5 percentage-point absolute diarization error tradeoff. Reported VoxConverse diarization error rates are 11.1% for Community-1 and 8.8% for Precision-2, and planned additions include voiceprint enrollment and cross-session speaker identification using embeddings already generated during diarization.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 6 | 265 | 57 | 33 | -89% |
| Real-time | 1 | 649 | 155 | 80 | -85% |
| Voice AI | 1 | 324 | 41 | 16 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.