How Lemon Slice built real-time generative video with Modal and Daily
Blog post from Modal
Lemon Slice uses Modal to operate AI character video products, evolving from a viral tool that generated speaking-character videos from images and text or audio into Lemon Slice Live, which enables real-time video conversations with AI characters. Rather than managing custom AWS or GCP infrastructure, the company deployed its 1-billion-parameter video model through two Python-based Modal Functions, allowing it to scale to 10,000 requests per hour with autoscaling GPU containers, reduced initialization times, and parallel model evaluation. For its live product, Modal launches separate containers for a Pipecat real-time processing server and GPU video inference, while Pipecat coordinates speech recognition from Deepgram, conversational responses from Grok, speech synthesis from ElevenLabs, and video generation. Direct TCP communication, regional co-location, and Daily’s WebRTC streaming infrastructure help reduce latency, producing video-and-audio responses in approximately three to six seconds.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.