A Practical Guide to Prompting Gemini 3.1 Flash TTS
Blog post from LiveKit
The guide provides detailed strategies for effectively using Gemini 3.1 Flash TTS, a speech synthesis model in beta, to generate natural and emotionally appropriate speech. It emphasizes the importance of structuring prompts with clear section labels and a "#### TRANSCRIPT" delimiter to prevent the model from reading stage directions aloud. The guide explains that Gemini interprets prompts as context, which requires precise direction to avoid confusion between narrative instructions and actual speech. Key practices include using commas instead of periods to maintain natural prosody, avoiding universal templates for emotional tags, and ensuring prompts begin with a synthesize-speech preamble. The guide also advises against using certain words like "quiet" or "flat" in style notes, as these can lead to monotone outputs, and recommends sticking to documented audio tags for the best results.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 1 | 5,932 | 1,046 | 223 | -2% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.