What it takes to build smarter voice agents: lessons from Retell and Super
Blog post from AssemblyAI
A panel featuring engineers from Retell and Super examined the practical challenges of deploying voice agents beyond demonstrations, emphasizing that production systems must handle noise, unreliable connections, repeat callers, and unpredictable conversational behavior. Both teams largely favor cascading architectures that separate speech-to-text, language-model, and text-to-speech components because they allow independent optimization, model switching, and provider redundancy, despite the convenience and potential latency benefits of end-to-end speech-to-speech systems. They target response times of roughly one to 1.5 seconds, balancing speed with a natural conversational pace, and assess performance through layered measures including word error rate, entity accuracy for important details such as names and account numbers, hard LLM test cases, call metrics, and human judgments of voice quality. Natural turn-taking remains difficult to quantify, requiring systems that avoid interrupting users, recognize genuine interruptions, and distinguish filler speech. The speakers identified persistent caller context as a key factor separating useful products from impressive demos, using structured profiles, extracted conversation variables, and CRM integrations to prevent users from repeating information. They also highlighted production monitoring, provider fallbacks, traffic routing, cost measurement, and real-world testing as essential operational practices, while anticipating more automated “loop engineering” systems that can identify failed calls and iteratively improve agents.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.