Home / Companies / Hume / Blog / Post Details
Content Deep Dive

Introducing the Hume Voice Controllability Leaderboard

Blog post from Hume

Post Details
Company
Date Published
Author
Hoon Shin, Shahbaz Mogal, and Alice Baird
Word Count
2,886
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Hume has introduced a Voice Controllability Leaderboard that evaluates 17 text-to-speech systems on their ability to create voices from descriptions, alter delivery through prompts or inline tags, and match voices to practical roles, using tens of thousands of human judgments alongside standardized acoustic measurements. The evaluation finds that no model leads all aspects of controllability: ElevenLabs v3 performs strongly in voice design, including accents, age, texture, multilingual capability, and role fit, while Google’s Gemini models lead tone, emotion, and instruction-following tests, and Inworld performs well in prosody, several accent categories, vocal bursts, podcasts, and therapist roles. Across models, gender control is highly reliable, but age, certain voice textures, authentic regional and multilingual accents, volume control, sarcasm, fear, disgust, and combinations or late placement of inline tags remain challenging. Results also show that naturalness and suitability for a requested role can diverge, and performance varies substantially by use case, with some lower-ranked models performing competitively in specialized categories. Hume argues that human listeners are essential for these assessments because language-model judging may not consistently reflect human experience or may be affected by evaluation data leakage.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.