Agora voice agent with AssemblyAI Universal-3 Pro Streaming
Blog post from AssemblyAI
Agora's voice agent integration with AssemblyAI Universal-3 Pro Streaming enables real-time transcription in Agora channels with minimal client-side changes. By utilizing a Python server as a silent observer, raw PCM audio from channel participants is streamed directly to AssemblyAI's WebSocket, achieving speaker-aware transcripts with a latency of 307ms P50. This setup leverages Agora's server-side bot capabilities to subscribe to participant audio and forward PCM streams to AssemblyAI, which processes them without the need for resampling. The integration provides significant improvements over Agora's built-in speech-to-text features, offering lower latency, better word error rates, and real-time speaker diarization across 99+ languages. The system architecture involves configuring the Agora channel for mono audio output at 16 kHz and setting up a websocket connection to stream participant audio frames to AssemblyAI, which in turn sends back transcript events for application logic or further processing.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.