Home / Companies / Modal / Blog / December 2024

December 2024 Summaries

6 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal announced upgrades and upcoming features aimed at improving AI and batch-computing workflows, including faster and more consistent beta memory snapshotting that can halve cold-start times for models such as Stable Diffusion, and planned async job queues supporting up to one million inputs. The platform now supports OpenID Connect for short-lived authentication tokens to access external services without long-lived credentials, as well as eStargz image compression to accelerate builds from registries such as ECR and Docker Hub. Recent client updates add detached app launches and expanded sandbox execution controls, while partnerships with Genmo and Chai Discovery simplify use and fine-tuning of open-source video-generation and molecular-prediction models. Modal also announced a multiyear AWS partnership enabling use of AWS Marketplace committed spend for GPUs, launched a YouTube channel with vLLM tutorials, and reported processing more than 40 billion inputs across over seven million apps in 2024.
Dec 28, 2024 367 words in the original blog post.
Modal has introduced a new inference-focused accelerator, the NVIDIA L40S GPU, priced at $1.95/hr, which offers substantial performance benefits over their current most popular accelerator, the NVIDIA A10 GPU. The L40S provides twice the on-device DDR6 random access memory of the A10, allowing users to run larger models on large inputs without a throughput-killing offload to CPU RAM. This results in a 40% speedup for memory-bound jobs and over a 100% speedup for compute-bound jobs using 16bit Tensor Cores. The L40S also outperforms the A10 in terms of streaming multiprocessor architecture, compute capability, and GPU RAM, with improved bandwidth and arithmetic capabilities. Modal users can now access this new accelerator through their platform, with $30/month in free compute available for sign-up.
Dec 19, 2024 466 words in the original blog post.
Fine-tuning adapts pretrained large language models to specific tasks or domains by updating their weights, often improving output quality while reducing prompt length, latency, and inference costs compared with general-purpose API models. It is suited to tasks such as structured-data generation, style control, and domain-specific classification, while retrieval-augmented generation may be preferable or complementary when external knowledge must be supplied at runtime. Fine-tuning can be performed through low-code services such as OpenAI and Predibase or configurable infrastructure including Modal, Google Colab, and AWS SageMaker, using frameworks such as Hugging Face Transformers, TRL, and Axolotl. The workflow involves selecting and testing an appropriate base model, preparing high-quality training and validation data with consistent prompt formats and tokenization, configuring and monitoring training, and applying efficiency techniques such as LoRA, QLoRA, quantization, and multi-GPU distributed training through DeepSpeed, FSDP, or Accelerate. Modal is presented as a serverless option for packaging training environments, accessing on-demand GPUs, and running distributed fine-tuning jobs for open-source models such as Llama and Mistral.
Dec 10, 2024 2,845 words in the original blog post.
This fall, a company went on an offsite to Ericeira, Portugal, where they organized an internal hackathon with around 30 employees. They were impressed by the results and wanted to share some of their favorite projects that used Modal in creative ways. The teams built various applications, including an AI agent app called Browserman, which navigated the internet to complete tasks, a real-time translation app called Glodal, and a workflow orchestrator called Waluigi. Another team developed a pytest plugin that parallelized test suites on Modal, shortening runtime from minutes to seconds. The hackathon showcased the potential of Modal to turbocharge developer productivity across diverse use cases in AI, arts, and software development.
Dec 09, 2024 611 words in the original blog post.
Modal partnered with a top soccer team, AFC Richmond, to help them process and analyze large amounts of tracking data from matches. The current system was not well-suited for this task due to limitations in scalability and cost-effectiveness. Modal's serverless batch processing on GPUs provided a flexible infrastructure that reduced costs by 50% and enabled the team to scale automatically based on data volume. The partnership also led to the development of a lightweight in-memory vector DB, allowing the coaching staff to make queries based on semantic similarity of embeddings generated during analysis. This solution enables AFC Richmond to gain valuable insights into player performance and make informed decisions for future matches.
Dec 04, 2024 525 words in the original blog post.
At Modal, a high-availability VPN proxy called vprox was built using Go and WireGuard. This proxy allows containers to funnel outbound traffic through static IPv4 addresses, ensuring consistent source IP addresses for outgoing internet traffic. The proxy uses SOCKS5 proxies initially but later switched to WireGuard, which provides better security and consistency. To configure networking for the proxy, a policy-based routing system was implemented, allowing traffic from one container to go through a designated VPN interface without affecting its neighbors. This setup is essential for multi-tenant Modal workers that run gVisor sandboxes on shared hosts. The proxy server uses sysctl settings to relax reverse path filtering and ensure reliable operation across Linux distributions. The vprox control plane was open-sourced, allowing developers to configure various aspects of the networking system, including IP discovery and client reconnection.
Dec 02, 2024 3,035 words in the original blog post.