Introducing Flux TTS: Conversation-Native Text to Speech for Real Time Voice Agents
Blog post from Deepgram
Deepgram has launched Flux TTS, a generally available text-to-speech model designed for real-time, multi-turn voice agents and offered free through September 12. Unlike narration-oriented TTS systems, Flux TTS is intended to retain conversational context across turns, adapt tone and pacing without SSML or detailed prompting, support interruptions by reporting what callers heard, and allow adjustments to speech characteristics while audio is being generated. Deepgram says the model was trained on conversational speech and uses a high-fidelity neural codec, interleaved text-and-audio generation, and a Mamba state-space architecture to combine expressive delivery, persistent context, and low latency, with first audio reported as low as 80 milliseconds. The company also reports benchmark advantages in word error rates, particularly for difficult production inputs such as account numbers, drug names, dates, currencies, and technical strings. Flux TTS can be deployed through cloud, self-hosted, or on-premises environments and is positioned for regulated industries and noisy settings such as restaurants. It integrates with Deepgramās Flux STT through a single API configuration, with planned shared state between the speech-to-text and text-to-speech models, while future features include additional languages, voice cloning, emotional controls, and expanded technical documentation.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.