Home / Companies / OpenRouter / Blog / Post Details
Content Deep Dive

New Audio APIs for Speech and Transcription

Blog post from OpenRouter

Post Details
Company
Date Published
Author
Jacky Liang
Word Count
623
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

OpenRouter introduces two specialized audio endpoints, /api/v1/audio/speech for text-to-speech and /api/v1/audio/transcriptions for speech-to-text, which deliver faster and more cost-efficient models tailored for specific audio tasks. These endpoints support generating speech from text using voices from providers like OpenAI, Google, and Mistral, as well as transcribing audio files with OpenAI Whisper, while maintaining consistent routing, billing, and key management across text, video, and image generation services. The choice between audio, speech, and transcription models depends on the desired balance of specialization, cost, and speed, with each model offering unique capabilities suitable for different use cases such as voice agents, reading text aloud, or creating meeting notes. Users can explore model functionalities in the Playground, selecting voices for speech models or uploading audio files for transcription, and each model page provides quickstart code examples in programming languages like Python and TypeScript. OpenRouter is continuously expanding its offerings with more providers and voices, and users are encouraged to suggest additional models through their Discord community.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 9,814 1,776 243 +42%
Real-time 1 6,790 1,736 269 -9%
Voice AI 1 4,562 308 52 +26%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.