Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Agora voice agent with AssemblyAI Universal-3 Pro Streaming

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
1,184
Company Posts That Month
44
Language
English
Hacker News Points
-
Post removed?
No
Summary

Agora's voice agent integration with AssemblyAI Universal-3 Pro Streaming enables real-time transcription in Agora channels with minimal client-side changes. By utilizing a Python server as a silent observer, raw PCM audio from channel participants is streamed directly to AssemblyAI's WebSocket, achieving speaker-aware transcripts with a latency of 307ms P50. This setup leverages Agora's server-side bot capabilities to subscribe to participant audio and forward PCM streams to AssemblyAI, which processes them without the need for resampling. The integration provides significant improvements over Agora's built-in speech-to-text features, offering lower latency, better word error rates, and real-time speaker diarization across 99+ languages. The system architecture involves configuring the Agora channel for mono audio output at 16 kHz and setting up a websocket connection to stream participant audio frames to AssemblyAI, which in turn sends back transcript events for application logic or further processing.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 22 6,296 1,346 246 -2%
Voice AI 8 2,379 221 38 -3%
LLM 4 5,932 1,046 223 -2%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.