Home / Companies / Gladia / Blog / Post Details
Content Deep Dive

Add speech-to-text to a Pipecat voice agent

Blog post from Gladia

Post Details
Company
Date Published
Author
Ani Ghazaryan
Word Count
2,895
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

Pipecat’s modular, frame-based architecture can integrate Gladia’s Solaria-1 speech-to-text service into real-time voice-agent pipelines alongside Daily WebRTC transport, Silero voice activity detection, LLMs, and text-to-speech services, allowing providers to be replaced without rewriting the overall workflow. The guide emphasizes an end-to-end conversational latency target below 500 ms, describing Solaria-1’s partial-transcript latency of under 103 ms and average response latency of roughly 300 ms, while noting that VAD settings such as the silence threshold strongly affect turn detection and responsiveness. It outlines required Python packages, 16 kHz PCM audio configuration, environment-variable-based API credentials, regional endpoint selection, WebSocket authentication and reconnection behavior, and production recovery practices such as retries, backoff, logging, and conversation-state checkpointing. It also discusses Solaria-1’s multilingual and code-switching support, sample-rate mismatch and network troubleshooting, pricing by streamed audio hour, data-training policies across plans, and the use of asynchronous post-processing for speaker diarization.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.