October 2026 Summaries
2 posts from Gradium
Filter
Month:
Year:
Post Summaries
Back to Blog
Gradium has added more than 1,000 AI voices to its catalog, expanding it from about 380 to 1,421 voices across English, Spanish, French, German, and Portuguese, with coverage for 28 accents and four use cases: customer service, narration, advertising and social media, and character performances. The release emphasizes improved naturalness and expressivity, with listener preference tests indicating that newly recommended customer-service voices generally outperformed previous flagship options; existing voices remain available to preserve current integrations. Voices were created through a voice-design model, screened from roughly 16,000 candidates by automated assessments for qualities such as audio quality, pacing, expressiveness, and accent, then evaluated by native listeners. The catalog includes 256 customer-service voices, 221 narration voices, 228 advertising and social-media voices, and 329 character voices, with Spanish receiving the largest language expansion. Users can filter voices in the studio by attributes including language, accent, gender, age, and use case, or use voice IDs in text-to-speech API requests. Future work will focus on more granular regional and city-level accents, with human listener evaluations continuing to determine which voices are released.
Oct 07, 2026
2,093 words in the original blog post.
Gradium’s latest text-to-speech model, now out of beta and the default for its API and studio, delivers its first audio chunk in roughly 50 milliseconds, positioning it as a low-latency option among frontier TTS systems. The company argues that its model improves on ElevenLabs v4 Turbo in latency, naturalness, expressivity, multilingual performance, and handling of difficult content such as phone numbers and contextual alphanumeric text. Although human conversations typically allow a 200–300 millisecond turn-taking gap, reducing TTS latency frees more of that budget for language-model reasoning, interruption handling, and network delays, potentially making voice agents feel more responsive, particularly during brief exchanges. The streaming model is positioned for real-time uses including support agents, coaching avatars, and telephone interactions, while pre-rendered applications such as audiobooks, narration, and IVR prompts can benefit from improved expressiveness. Existing API users do not need to make changes, and developers may provide longer opening phrases for additional context without sacrificing startup speed.
Oct 01, 2026
474 words in the original blog post.