Home / Companies / Coval / Blog / Post Details
Content Deep Dive

The State of Voice AI Instruction Following in 2026: A Conversation with Kwindla from Pipecat and Zach from Ultravox

Blog post from Coval

Post Details
Company
Date Published
Author
Brooke Hopkins
Word Count
1,755
Company Posts That Month
15
Language
English
Hacker News Points
-
Post removed?
No
Summary

In a detailed exploration of the current state of voice AI, experts Kwindla Hultman Kramer and Zach Koch discuss the challenges and advancements in the field, particularly focusing on the issues around evaluation and model deployment. Despite the existence of advanced models like GPT-5, the industry still relies on older models such as GPT-4o due to their optimal balance of intelligence and latency, and the complexities involved in switching models. The conversation highlights the difficulty of benchmarking instruction following due to varying performance across different models and the lack of adequate representation of long, multi-turn interactions in training data. Key challenges identified include the absence of benchmarks for back-channeling, prosody matching, and subtle timing issues that affect conversational naturalness. The discussion also touches on the emerging trend of using multi-model architectures, which, while promising, add layers of complexity to evaluation. The importance of feedback from real-world deployments is emphasized as a crucial factor for model improvement in this rapidly evolving space.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.