Introducing EVI 3: the world’s most realistic and instructible speech-to-speech foundation model
Blog post from Hume
Hume has introduced its third-generation speech-language model, EVI 3, which brings more expressiveness, realism, and emotional understanding to voice AI experiences. EVI 3 can be fully personalized with any voice and personality created by a user prompt, allowing for instant generation of new voices and personalities. This is achieved through Hume's latest research on speech-language models, which developed methods to capture the full range of human voices and speaking styles in one model. EVI 3 has been evaluated against other leading voice-to-voice AI models, including GPT-4o, Gemini, and Sesame, with favorable results in areas such as emotional tone modulation, expressiveness, and naturalness. The model is capable of delivering voice responses in under 300ms on state-of-the-art hardware and is currently available through a live demo and iOS app, with API access planned for release in the coming weeks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 4 | 4,558 | 674 | 207 | -8% |
| Voice AI | 4 | 1,094 | 163 | 44 | +63% |
| Real-time | 3 | 4,099 | 1,129 | 265 | -46% |
| AI Model Fine-tuning | 1 | 790 | 187 | 78 | -8% |
| Reinforcement learning | 1 | 175 | 93 | 31 | -18% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.