AI dubbing API: Dub audio and video with an API at scale
Blog post from ElevenLabs
An AI dubbing API automates localization of audio and video by transcribing source media, translating dialogue, cloning speakers’ voices, synthesizing speech in target languages, and aligning the output with original timing. ElevenLabs’ Dubbing v2 uses an audio-to-audio approach intended to preserve performance characteristics such as tone, emotion, pitch, and delivery more effectively than cascaded speech-to-text, translation, and text-to-speech pipelines, while supporting sync-aware translation, multiple speakers, regional language variants, and more than 90 languages. Developers create a project from a media file or URL, wait for transcription to finish, add target languages with configurable voice-cloning strength, and retrieve completed dubbed audio through signed download links. Enterprise users can edit source and translated transcript segments, supply their own transcripts or translations, and regenerate only modified portions, with free regeneration up to the source media’s duration. The API is positioned for creator tools, training, streaming, marketing, and education workflows; Dubbing v2 handles prerecorded media and starts at $2.20 per source minute, while the lower-cost Dubbing v1 uses a cascaded pipeline at $0.33 per minute.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.