Add speech-to-text to a VAPI voice agent
Blog post from Gladia
The guide explains how to replace Vapi’s default speech-to-text provider with Gladia’s Solaria-1 real-time transcription model through Vapi’s custom WebSocket transcriber interface, arguing that STT accuracy and latency can significantly affect downstream LLM tool calls, routing, and conversational flow. It describes creating a per-session Gladia WebSocket URL through a live API request, supplying it to Vapi’s assistant configuration, and using recommended telephony audio settings, with no required LLM or text-to-speech changes. Gladia reports partial transcripts in under 103 milliseconds and final transcripts averaging about 270 milliseconds, positioning this within an approximately one-second end-to-end voice-agent latency budget. The guide covers endpointing adjustments, partial-transcript processing, automatic language detection and code-switching for more than 100 languages, and custom vocabulary for specialized terms, while advising teams to benchmark performance on their own audio rather than relying solely on vendor figures. It also notes that real-time speaker diarization is unavailable and should be handled after calls, outlines reliability, compliance, data-training, concurrency, and fallback considerations, and compares stated pricing and capabilities with Deepgram and AssemblyAI.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.