The Future of Speech-to-Speech AI: Inside Gradium and Kyutai's Approach to Full Duplex Conversation
Blog post from Coval
In the latest episode of "Conversations in Conversational AI," Neil Zeghidour, CEO and co-founder of Gradium and co-founder of Kyutai, discusses the innovative developments in speech-to-speech models that are reshaping the future of voice AI. Zeghidour's journey began with the establishment of Kyutai as a nonprofit research lab, emphasizing the importance of risk-taking in breakthrough innovations, which eventually led to the creation of Gradium for commercial product development. Key advancements include the Moshi project, which introduced full duplex conversations, allowing simultaneous speaking and listening without traditional turn-taking constraints, achieved through audio language models instead of diffusion models. Despite these advancements, challenges remain due to the intelligence gap between speech-to-speech and text models, largely due to differences in training data and inherent distractions in audio data. While cascaded systems currently dominate, offering modularity and steerability, Zeghidour predicts a shift towards more natural, expressive interactions that capture the complexities of human conversation. The episode also highlights the potential of miniaturized models like Pocket TTS for efficient on-device processing and explores the future applications of voice AI in robotics and spatial audio environments, where current models struggle to adapt.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.