How sync. uses Modal to lipsync 100 hours of video a day
Blog post from Modal
Sync, a research lab founded by the team behind Wav2Lip, develops foundational AI models for manipulating humans in video, including a zero-shot lip-syncing system that can preserve a speaker’s style across translated languages. After gaining attention with a viral Hindi-language video demo, the five-person team needed to move beyond Google Colab and infrastructure tools such as AWS Lambda, which created deployment, scaling, and GPU-support challenges, while Replicate’s container rebuild process slowed updates. Using Modal’s startup credits and developer workflow, Sync was able to deploy changes rapidly, relying on Modal for autoscaling and GPU infrastructure rather than MLops management. The company reports deploying up to 95 times daily, releasing 10 major production model variants and roughly 1,000 intermediate iterations in a year. Its production workflow processes more than 100 hours of video daily by splitting long videos into scenes, running parallel face detection and translation on T4 GPUs, applying its lip-sync model on A100 GPUs, and stitching the results together. Sync plans to expand beyond translation and lip-syncing into AI-powered emotion editing, pose adjustment, and changes to subjects’ physical characteristics.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Serverless | 1 | 1,628 | 326 | 111 | +97% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.