New Audio APIs for Speech and Transcription
Blog post from OpenRouter
OpenRouter introduces two specialized audio endpoints, /api/v1/audio/speech for text-to-speech and /api/v1/audio/transcriptions for speech-to-text, which deliver faster and more cost-efficient models tailored for specific audio tasks. These endpoints support generating speech from text using voices from providers like OpenAI, Google, and Mistral, as well as transcribing audio files with OpenAI Whisper, while maintaining consistent routing, billing, and key management across text, video, and image generation services. The choice between audio, speech, and transcription models depends on the desired balance of specialization, cost, and speed, with each model offering unique capabilities suitable for different use cases such as voice agents, reading text aloud, or creating meeting notes. Users can explore model functionalities in the Playground, selecting voices for speech models or uploading audio files for transcription, and each model page provides quickstart code examples in programming languages like Python and TypeScript. OpenRouter is continuously expanding its offerings with more providers and voices, and users are encouraged to suggest additional models through their Discord community.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.