Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

What it takes to build smarter voice agents: lessons from Retell and Super

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
2,040
Company Posts That Month
21
Language
English
Hacker News Points
-
Post removed?
No
Summary

A panel featuring engineers from Retell and Super examined the practical challenges of deploying voice agents beyond demonstrations, emphasizing that production systems must handle noise, unreliable connections, repeat callers, and unpredictable conversational behavior. Both teams largely favor cascading architectures that separate speech-to-text, language-model, and text-to-speech components because they allow independent optimization, model switching, and provider redundancy, despite the convenience and potential latency benefits of end-to-end speech-to-speech systems. They target response times of roughly one to 1.5 seconds, balancing speed with a natural conversational pace, and assess performance through layered measures including word error rate, entity accuracy for important details such as names and account numbers, hard LLM test cases, call metrics, and human judgments of voice quality. Natural turn-taking remains difficult to quantify, requiring systems that avoid interrupting users, recognize genuine interruptions, and distinguish filler speech. The speakers identified persistent caller context as a key factor separating useful products from impressive demos, using structured profiles, extracted conversation variables, and CRM integrations to prevent users from repeating information. They also highlighted production monitoring, provider fallbacks, traffic routing, cost measurement, and real-world testing as essential operational practices, while anticipating more automated “loop engineering” systems that can identify failed calls and iteratively improve agents.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.