January 2025 Summaries
7 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Modal has introduced memory snapshots to reduce serverless GPU container cold starts by capturing a container’s filesystem changes and full process state just before it accepts requests, then restoring that state later without rerunning costly initialization work such as Python imports. Built on gVisor’s checkpoint/restore capabilities, the approach is especially effective for workloads like PyTorch, whose imports trigger tens of thousands of filesystem-related system calls; Modal combines restored memory mappings with prioritized background page loading and aggressive caching to accelerate startup. Tests show snapshot restores averaging roughly 2.5 times faster than standard starts, reducing a Stable Diffusion function from about 13 seconds to 3.5 seconds and a simple PyTorch import from roughly five seconds to near one second. Snapshots are created on demand and may be specific to compatible worker hardware, CPU features, NVIDIA driver versions, and runtime versions, so Modal manages their lifecycle and falls back to ordinary startup if restoration fails. The feature requires applications to tolerate paused and reused process state, cannot preserve live network connections or GPU state, and therefore supports lifecycle hooks that snapshot CPU-side model setup while initializing GPU memory after restoration.
Jan 28, 2025
2,020 words in the original blog post.
The MTEB leaderboard is a comprehensive benchmark that evaluates the performance of embedding models across various tasks, providing a standardized way to compare different models. While high ranking on the leaderboard doesn't guarantee the best fit for a specific use case, considering factors such as task-specific performance, computational requirements, and domain relevance can help make an informed decision. Top models currently on the MTEB leaderboard include generalist embedding models like NV-Embed-v2, Nomic-Embed-Text-v1.5, and bge-en-icl, which have been fine-tuned for specific tasks or domains such as medicine, finance, law, code, math, Japanese, Korean, Chinese, French, Arabic, among others. Domain-specific embedding models can offer superior performance for specialized applications, making it essential to explore these models alongside top performers on the leaderboard to find the best fit for a particular use case.
Jan 27, 2025
701 words in the original blog post.
Axolotl is a wrapper for lower-level Hugging Face libraries that simplifies the fine-tuning process of large language models, offering granular control while being easier to use. It comes with built-in default values and optimizations, including sample packing, which can improve training efficiency. Axolotl allows users to train open weights models like LLaMA 3/LLaMA 3.1, Pythia, and Falcon on their own data without needing to implement the fine-tuning process from scratch. Unsloth is a framework designed to dramatically improve the speed and efficiency of LLM fine-tuning, allowing users to fine-tune Llama 3.1, Mistral, Phi & Gemma LLMs up to 2-5x faster with 80% less memory usage compared to FA2. Torchtune is a PyTorch-native library for easily fine-tuning LLMs, offering a lean and extensible design that's just pure PyTorch, with excellent interoperability with popular libraries across the PyTorch ecosystem. The choice between these tools ultimately depends on specific requirements, hardware constraints, and level of expertise, with Axolotl being recommended as a good starting point for beginners.
Jan 27, 2025
614 words in the original blog post.
NVIDIA's A10, A100, and H100 GPUs offer varying performance-to-cost ratios, making them suitable for different machine learning tasks. The H100 is ideal for large language model workloads with high precision requirements, while the A100 is a versatile GPU suitable for larger models with moderate precision needs. In contrast, the A10 and L4 GPUs are more cost-effective options for smaller models or inference tasks. When selecting a GPU, consider factors such as task type, model size, memory requirements, budget, and performance needs to choose the best fit for your specific use case.
Jan 27, 2025
844 words in the original blog post.
Modal has announced the general availability of Sandboxes, an isolated execution environment designed to safely run untrusted code from LLMs, users, or third parties without exposing a customer’s Modal workspace or infrastructure. Sandboxes support arbitrary languages, runtime-configurable dependencies through Modal’s Image API, filesystem snapshots for persistent and parallelized workloads, and controls for networking, ports, and file access. Built on the same infrastructure as Modal Functions, they offer fast cold starts, GPU availability, global region selection, and integration with existing Modal applications. Customers including SWE-bench, Quora’s Poe platform, Codegen, and Relevance AI use Sandboxes for parallel agent evaluations, secure interactive coding, AI-driven code refactoring, and automated workflows, with Modal reporting tests of creation throughput up to 1,000 Sandboxes per second.
Jan 21, 2025
1,070 words in the original blog post.
Modal has introduced NVIDIA L40S GPUs at $1.95 per hour, offering 48GB of memory and substantially improved performance over A10 GPUs for larger models, longer contexts, and memory- or compute-intensive workloads. The platform also added proxy authentication tokens to restrict access to web endpoints, a Filesystem API and disk snapshotting for Sandboxes to support interactive file operations, state restoration, branching, and reduced cold starts. Other client updates include automatic `.dockerignore` handling, file-pattern loading, volume renaming, sandbox file watching and larger write payloads, environment selection in `App.run`, and configurable Docker images for VSCode launches. Modal additionally announced SOC 2 Type 2 certification, published a GPU glossary, highlighted resources for protein modeling, OCR, and diffusion-model deployment, and noted its biotech community dinners.
Jan 21, 2025
545 words in the original blog post.
As a cloud compute platform, Modal is committed to customer security and privacy as top priorities. Our product is secure by design, and we take measures from development to deployment to mitigate risk and earn the trust of our users. We have achieved SOC 2 Type I compliance last year and announced support for HIPAA compliant workloads earlier this year. Recently, we completed a more rigorous SOC 2 Type II audit with no deviations found, demonstrating our commitment to continually improving our security posture. As our customer base grows, we will renew our SOC 2 audits annually alongside product and process improvements, providing stronger reassurances around the security of our platform.
Jan 02, 2025
216 words in the original blog post.