New Audio APIs for Speech and Transcription
Blog post from OpenRouter
OpenRouter has introduced two dedicated audio endpoints for enhanced speech and transcription capabilities: /api/v1/audio/speech for text-to-speech and /api/v1/audio/transcriptions for speech-to-text. These endpoints offer specialized models that are faster and more cost-efficient for specific audio tasks compared to general audio models. Users can generate speech using voices from OpenAI, Google, or Mistral and transcribe audio files with OpenAI Whisper, maintaining consistent routing, billing, and key management across different media types. The choice between audio, speech, and transcription models involves a trade-off between specialization, cost, and speed, with each model being optimized for particular use cases such as voice agents, reading text aloud, or transcribing meeting notes. The platform provides tools like a Playground for experimenting with models and quickstart code samples for integration, supporting a variety of audio formats and allowing for customization such as tone control. OpenRouter plans to expand its offerings with more providers and voices, inviting user feedback for future developments.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.