Introducing the Hume Voice Controllability Leaderboard
Blog post from Hume
Hume has introduced a Voice Controllability Leaderboard that evaluates 17 text-to-speech systems on their ability to create voices from descriptions, alter delivery through prompts or inline tags, and match voices to practical roles, using tens of thousands of human judgments alongside standardized acoustic measurements. The evaluation finds that no model leads all aspects of controllability: ElevenLabs v3 performs strongly in voice design, including accents, age, texture, multilingual capability, and role fit, while Google’s Gemini models lead tone, emotion, and instruction-following tests, and Inworld performs well in prosody, several accent categories, vocal bursts, podcasts, and therapist roles. Across models, gender control is highly reliable, but age, certain voice textures, authentic regional and multilingual accents, volume control, sarcasm, fear, disgust, and combinations or late placement of inline tags remain challenging. Results also show that naturalness and suitability for a requested role can diverge, and performance varies substantially by use case, with some lower-ranked models performing competitively in specialized categories. Hume argues that human listeners are essential for these assessments because language-model judging may not consistently reflect human experience or may be affected by evaluation data leakage.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.