August 2026 Summaries
3 posts from Gradium
Filter
Month:
Year:
Post Summaries
Back to Blog
Gradium has released a new text-to-speech model as its default offering, emphasizing accurate real-time pronunciation of structured information such as phone numbers, email addresses, financial amounts, dates, account names, and reference codes without user-side text normalization. In its August 2026 benchmark of 500 difficult sentences across English, German, French, Spanish, and Portuguese, Gradium reports an 81.0% human-evaluated pass rate, compared with 75.1% for Cartesia Sonic 3.6 and lower scores for several other tested models, alongside a 216 ms median time to first audio and a 30 ms interquartile latency spread. The company says its production and studio paths use the same unmodified text input, aiming to avoid hidden rewriting and hallucinated content, and it has open-sourced the evaluation dataset on Hugging Face. Gradium invites users to submit additional failure cases for future evaluations, offers credits for complete reports, and says further work will target latency, naturalness, accuracy, and multilingual coverage; the model is available through its API and Studio while retaining compatibility with existing and custom voices.
Aug 31, 2026
1,352 words in the original blog post.
AudioStack has integrated Gradium as a voice provider, making its multilingual voices and regional accents available through the platform’s audio-production tools. AudioStack automates script generation, voice selection, music, sound design, mixing, and mastering, enabling users to create large volumes of localized advertising, publishing, and other audio content through an API or console. The partnership emphasizes Gradium’s regional voice depth, including Australian English, French Canadian, Dublin English, and Bavarian German, which AudioStack says can help customers produce more contextually relevant content for local markets. Both companies describe the collaboration as supporting AudioStack’s broader strategy of expanding its network of voice, music, and sound providers while allowing media organizations to create broadcast-ready audio at scale.
Aug 19, 2026
461 words in the original blog post.
Gradium describes voice selection as a structured casting process that combines customer-defined requirements with large-scale generated candidate pools and listener evaluation. Voice targets are specified by locale, demographic characteristics, vocal traits, intended use case, and delivery style, with use case treated as especially influential because listener preferences vary between contexts such as customer support and narration. The company recommends defining the actual listener, creating detailed personas, identifying undesirable brand traits, translating descriptive language into acoustic instructions with native-speaker input, and using realistic evaluation scripts. Flagship voices are selected through diverse generation, crowd-based keeper tests, head-to-head ELO ranking against existing catalog voices, statistical confidence thresholds, and native-speaker checks for accent and pronunciation. The approach emphasizes that voice descriptors and preferences do not transfer reliably across languages or cultures, requiring locale-specific prompts and evaluation, while accent assessment depends on scripts that expose distinguishing sounds and validation by native listeners. Gradium reports a catalog of more than 360 voices across English, French, German, Portuguese, and Spanish, and plans to expand locales, use cases, voice discovery features, and access to its voice-design model.
Aug 04, 2026
2,362 words in the original blog post.