Home / Companies / Deepgram / Blog / Post Details
Content Deep Dive

Build A Voice Agent With Pipecat And Deepgram Flux STT And Flux TTS

Blog post from Deepgram

Post Details
Company
Date Published
Author
Jose Nicholas Francisco
Word Count
2,477
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

Deepgram’s guide explains how to build a real-time Pipecat voice agent using Flux STT for conversational speech recognition and turn detection, Flux TTS for streaming speech synthesis, and an LLM such as OpenAI between them. Unlike conventional pipelines that separately use voice activity detection and turn-analysis components, Flux STT combines transcription and model-native turn detection, reducing configuration and potentially lowering response latency, while Pipecat provides orchestration across audio transport, speech services, and language models. The setup uses the Pipecat CLI to scaffold a Daily WebRTC-based cascade agent, requires Deepgram, OpenAI, and Daily API keys, and supports configuration of Flux STT events such as start, end, eager end, and resumed turns. Flux TTS uses the dedicated DeepgramFluxTTSService, with voices such as the default flux-alexis-en, and streams audio through Deepgram’s Speak v2 endpoint. The guide also covers local testing, interruption behavior, alternative WebRTC and telephony transports, metrics for monitoring latency and turn quality, and cloud, self-hosted, on-premises, VPC, and SageMaker deployment options for organizations with scalability or compliance requirements.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.