Introducing the Hume Voice Replication Leaderboard
Blog post from Hume
Hume has introduced the Voice Replication Leaderboard, part of its Real World VoiceEQ Benchmark, to assess how effectively 11 text-to-speech models clone a speaker’s identity while maintaining audio quality and natural delivery. Models replicated 25 voices spanning standard speech, emotional recordings, and native and non-native English accents using identical prompts, then received human ratings for same-speaker similarity, quality, and naturalness alongside an objective speaker-embedding similarity score. OpenBMB VoxCPM2 led human-rated speaker similarity, LongCat-AudioDiT-3.5B achieved the highest objective similarity, Cartesia sonic-3.6-beta led naturalness, and Inworld TTS-2 led audio quality, demonstrating that no system was best in every category. Results also showed substantial differences by reference type, with expressive speech posing particular difficulty for some models, while quality ratings were generally high and less variable than identity preservation. Hume argues that developers and product teams should evaluate models across separate dimensions and with deployment-relevant voices, accents, emotions, and conditions, combining human judgment with objective metrics rather than relying on a single overall score.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 3 | 265 | 57 | 33 | -89% |
| Voice AI | 2 | 324 | 41 | 16 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.