Updates for developers building with voice
Blog post from OpenAI
AI audio capabilities have recently been enhanced with the release of new model snapshots aimed at improving the reliability and quality of audio agents in various workflows, such as transcription, text-to-speech, and speech-to-speech. These updates, which include models like gpt-4o-mini-transcribe-2025-12-15 and gpt-realtime-mini-2025-12-15, offer significant improvements in accuracy, natural voice output, and decreased error rates, particularly in noisy environments. The new models also demonstrate enhanced performance in instruction following and tool calling, making them suitable for real-time applications where cost and latency are critical. Notably, the updates support Custom Voices, allowing organizations to create unique brand voices with better natural tones and dialect accuracy. The advancements are designed to address common challenges in voice applications, such as handling long conversations and edge cases, by reducing errors and hallucinations and ensuring consistent tool use. Developers are encouraged to adopt these new snapshots for enhanced performance at no additional cost.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.