Image Generation MCP: Images, Video & Lip Sync with Fish Audio
Blog post from Fish Audio
Fish Audio’s MCP server has expanded from audio functions to include image generation and editing, video generation, lip sync, and talking-avatar creation through a single authenticated connection for compatible AI agents such as Claude Code, Cursor, Codex, Windsurf, and Devin. Users can access a changing selection of third-party media models without managing separate provider API keys, while sharing one Fish Audio credit balance and reusing generated files directly across subsequent tasks. The platform supports workflows such as generating images, animating them into videos, creating speech with text-to-speech, and passing that audio into avatar or lip-sync models without manual downloads and uploads. Media jobs follow a quote, confirmation, and generation process designed to prevent unapproved charges, while failed jobs are automatically refunded and duplicate requests are protected from repeat billing; text-to-speech is an exception, billed immediately by text length. Image generation, editing, and upscaling are available on the free plan, whereas video, lip sync, and avatar capabilities require a paid Plus plan or above, with costs varying by model and output settings.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.