February 2024 Summaries
4 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Modal has introduced WebSocket support, enabling persistent, bidirectional, low-latency communication between clients and server-side functions for real-time applications. Developers can use standard libraries such as FastAPI to create WebSocket endpoints and deploy them through Modal without specialized connection setup, although functions should allow concurrent inputs when workloads are not CPU- or GPU-bound to avoid launching a separate container per connection. Modal automatically scales WebSocket handlers and related functions according to demand, supporting applications such as multi-user real-time speech transcription. Suggested uses include streaming outputs for audio, text-to-speech, and image-generation systems; delivering progress updates for long-running tasks such as large-language-model analysis; and hosting WebSocket-dependent open-source tools including ComfyUI, Streamlit, and Gradio for interactive machine-learning interfaces.
Feb 27, 2024
562 words in the original blog post.
Suno, a music generation app, uses Modal to scale its inference and batch pre-processing capabilities, allowing it to bring a state-of-the-art model to market four months early. By avoiding the need to manage their own clusters and diverting engineering resources, Suno was able to focus on developing its core product. With Modal's auto-scaling feature, Suno can efficiently handle increased demand during holidays, saving time and financial commitments. The partnership with Microsoft to integrate song generation capabilities into Copilot further highlights Suno's innovative approach to AI-powered music creation.
Feb 21, 2024
509 words in the original blog post.
Modal’s 2024 updates add interactive container commands through `modal container exec`, enabling users to inspect running jobs, monitor processes, and troubleshoot environments directly. The platform now supports WebSockets for Modal functions, allowing applications built with frameworks such as Streamlit to be hosted more easily, and it has introduced NVIDIA H100 GPU availability. Client improvements include decorator-based build lifecycle hooks, automatic background commits for Volumes, expanded Volume CLI operations, enhanced Sandbox management, image-specific import declarations, private Docker registry authentication, and dictionary clearing, with a detailed changelog available for ongoing updates. Modal also highlights new example projects, including Turbo.art, rapid English Wikipedia embedding, and LLM fine-tuning workflows, while inviting users to share their own work through its community Slack.
Feb 15, 2024
219 words in the original blog post.
Modal has introduced access to NVIDIA H100 GPUs for its users, positioning the hardware as a high-performance option for machine learning workloads such as latency-sensitive LLM inference and large-scale model fine-tuning. NVIDIA benchmarks cited in the announcement indicate that H100s can provide up to four times faster training and up to 30 times faster inference than A100 GPUs for large language models. Modal’s H100 instances provide 80 GB of memory per GPU and support configurations of up to eight GPUs connected through NVLink, priced at $7.65 per GPU hour. The platform emphasizes usage-based billing and autoscaling as ways to improve utilization and potentially reduce costs compared with fixed GPU reservations, while users can request an H100 by specifying it as the GPU type in a remote function configuration.
Feb 06, 2024
312 words in the original blog post.