The State of Voice AI Instruction Following in 2026: A Conversation with Kwindla from Pipecat and Zach from Ultravox
Blog post from Coval
In a detailed exploration of the current state of voice AI, experts Kwindla Hultman Kramer and Zach Koch discuss the challenges and advancements in the field, particularly focusing on the issues around evaluation and model deployment. Despite the existence of advanced models like GPT-5, the industry still relies on older models such as GPT-4o due to their optimal balance of intelligence and latency, and the complexities involved in switching models. The conversation highlights the difficulty of benchmarking instruction following due to varying performance across different models and the lack of adequate representation of long, multi-turn interactions in training data. Key challenges identified include the absence of benchmarks for back-channeling, prosody matching, and subtle timing issues that affect conversational naturalness. The discussion also touches on the emerging trend of using multi-model architectures, which, while promising, add layers of complexity to evaluation. The importance of feedback from real-world deployments is emphasized as a crucial factor for model improvement in this rapidly evolving space.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.