Making voice agents sound human with expressive mode
Blog post from LiveKit
LiveKit has introduced expressive mode for its Agents framework to help voice AI adapt its tone, pacing, emotions, and nonverbal sounds to a user’s conversational context, addressing what it describes as a “feeling skills gap” that can reduce customer satisfaction in emotionally sensitive interactions. Enabled through a single configuration flag, the feature prompts an LLM to generate inline expressive markup, translates that markup into the dialect required by supported text-to-speech providers, cleans it from user-facing transcripts, and batches speech into larger chunks to preserve emotional consistency without meaningful added latency. It currently supports markup-capable LiveKit Inference models from Fish Audio, Inworld, Cartesia, and xAI, while allowing developers to selectively constrain delivery settings or customize the injected instructions. Expressive metadata is also made available to frontends through an lk.expression transcription attribute and a UI hook that converts provider-specific labels into normalized moods and colors for visualizers or indicators. LiveKit plans to add live delivery controls, broader provider support, and richer frontend expression data in future releases.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.