TTS with emotion: emphasis and control in live conversations
Blog post from Deepgram
Live voice agents must manage emotional TTS without the manual review available in traditional voiceover workflows, using either inline markup generated by the LLM, session-level synthesis parameters, or cross-turn model state. Using Deepgram’s Flux TTS as an example, the text explains that inline markup enables turn-specific expression but adds token cost, buffering requirements, and risks from malformed tags, while Flux’s connection-level expressivity setting ranges from calm to animated and provides consistent delivery without added token or latency costs but remains fixed for a WebSocket session and is labeled beta. Flux also preserves prosodic context across turns within an active connection, although this state cannot be directly adjusted and resets only when a new connection is opened. Because emphasis involves linked changes in pitch, duration, and energy, emotion settings can affect voices differently and cannot reliably solve pronunciation issues, which may instead require tools such as Aura-2’s IPA lexicon controls. The guidance recommends default expressivity and voice selection for accuracy-sensitive healthcare, financial, and IVR readbacks, while allowing carefully auditioned non-default settings or markup for lower-risk conversational, sales, and reminder use cases. It advises testing exact voices, scripts, number-heavy content, reconnection behavior, latency, and model updates before deployment, since stronger expressivity settings can increase repetitions, omitted words, added words, and pronunciation errors.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.