Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Streaming real-time text to speech with XTTS V2

Blog post from Baseten

Post Details
Company
Date Published
Author
Het Trivedi, Philip Kiely
Word Count
1,318
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

A streaming endpoint for XTTS V2, a state-of-the-art open-source text-to-speech model with voice cloning capabilities, can be deployed to power an entire new class of AI applications. The streaming endpoint has a round-trip time to first chunk of as little as 200 milliseconds and delivers near real-time audio playback for a given text input. XTTS V2 is natively capable of streaming and can generate speech in 17 languages, with the ability to support over a dozen languages. A model server implemented in Truss enables fast inference times, and deploying the streaming endpoint requires setting GPU resources in config.yaml and running `truss push` to create a development deployment on Baseten. Consuming the model output depends on the application, but can be demonstrated with a quick Python script that streams the audio with FFmpeg.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 21 2,334 631 194 -8%
LLM 3 3,398 379 136 +44%
Voice AI 2 188 74 20 +24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.