Home / Companies / Hume / Blog / September 2026

September 2026 Summaries

2 posts from Hume

Filter
Month: Year:
Post Summaries Back to Blog
Hume has introduced a Voice Controllability Leaderboard that evaluates 17 text-to-speech systems on their ability to create voices from descriptions, alter delivery through prompts or inline tags, and match voices to practical roles, using tens of thousands of human judgments alongside standardized acoustic measurements. The evaluation finds that no model leads all aspects of controllability: ElevenLabs v3 performs strongly in voice design, including accents, age, texture, multilingual capability, and role fit, while Google’s Gemini models lead tone, emotion, and instruction-following tests, and Inworld performs well in prosody, several accent categories, vocal bursts, podcasts, and therapist roles. Across models, gender control is highly reliable, but age, certain voice textures, authentic regional and multilingual accents, volume control, sarcasm, fear, disgust, and combinations or late placement of inline tags remain challenging. Results also show that naturalness and suitability for a requested role can diverge, and performance varies substantially by use case, with some lower-ranked models performing competitively in specialized categories. Hume argues that human listeners are essential for these assessments because language-model judging may not consistently reflect human experience or may be affected by evaluation data leakage.
Sep 17, 2026 2,886 words in the original blog post.
Hume has introduced the Voice Replication Leaderboard, part of its Real World VoiceEQ Benchmark, to assess how effectively 11 text-to-speech models clone a speaker’s identity while maintaining audio quality and natural delivery. Models replicated 25 voices spanning standard speech, emotional recordings, and native and non-native English accents using identical prompts, then received human ratings for same-speaker similarity, quality, and naturalness alongside an objective speaker-embedding similarity score. OpenBMB VoxCPM2 led human-rated speaker similarity, LongCat-AudioDiT-3.5B achieved the highest objective similarity, Cartesia sonic-3.6-beta led naturalness, and Inworld TTS-2 led audio quality, demonstrating that no system was best in every category. Results also showed substantial differences by reference type, with expressive speech posing particular difficulty for some models, while quality ratings were generally high and less variable than identity preservation. Hume argues that developers and product teams should evaluate models across separate dimensions and with deployment-relevant voices, accents, emotions, and conditions, combining human judgment with objective metrics rather than relying on a single overall score.
Sep 10, 2026 1,545 words in the original blog post.