May 2025 Summaries
8 posts from Lambda
Filter
Month:
Year:
Post Summaries
Back to Blog
Qwen3-32B, a dense model developed by Alibaba, is now available on Lambda's Inference API, offering advanced capabilities such as hybrid reasoning, multilingual support, and agentic capacities. With its STEM and logical reasoning proficiency, Qwen3-32B can process complex tasks that require human intervention, including coding, math, and logic. The model features two problem-solving modes, thinking mode and non-thinking mode, allowing developers to switch between sequential processing and instant responses. It also supports 119 languages and dialects, creative writing, role-playing, instruction following, and multi-turn dialogue. Qwen3-32B can execute agentic actions, enabling developers to call their tools of choice and modify Model Context Protocol configuration files. The model was pre-trained with 36 trillion tokens from online sources and PDF documents, delivering similar performance to the Qwen2.5 base models while operating more efficiently due to improved pre-training processes. Qwen3-32B is available on Lambda's Inference API for $0.10 per million input tokens and $0.30 per million output tokens, with no rate limits or cost-efficient pricing.
May 29, 2025
815 words in the original blog post.
In a discussion hosted by Emma Brooks of DCD, Phil Lawson-Shanks of Aligned Data Centres and Ken Patchett of Lambda examined the industry shift from traditional air cooling to liquid and hybrid systems in data centers, driven by the increasing power demands of AI workloads. They discussed how next-generation chips, like NVIDIA's Blackwell, necessitate new cooling strategies due to their high power densities, emphasizing the importance of adaptability and scalability in data center design. The conversation highlighted the transformational opportunities liquid cooling offers for efficiency and sustainability, as well as the need for updated operational procedures and skill sets to manage these technologies. Both experts noted the rapid evolution of data center infrastructure to accommodate future AI capabilities, stressing the importance of collaboration between data center operators and hyperscalers to support the ongoing technological renaissance.
May 21, 2025
8,547 words in the original blog post.
MLflow is an open-source platform designed to streamline and manage the machine learning lifecycle, from experimentation to deployment. Lambda offers high-performance GPU instances that are perfect for training and deploying machine learning models. To implement MLflow on Lambda Cloud's on-demand instances, you need a Lambda Cloud account, SSH key set up for accessing remote instances, and some familiarity with MLflow's features and functions at a high-level. The platform simplifies the model lifecycle deployment process by combining Lambda's infrastructure with MLflow's experiment tracking capabilities. You can track (logs parameters and results), projects (packages code), models (manages and deploys models), and registry (centrally stores models) using MLflow. To set up your environment, you need to launch an on-demand instance with appropriate GPU resources, select a Linux distribution, and configure network and firewall settings. You can then install Python, dependencies, and MLflow, set up storage for artifacts, start the MLflow tracking server, and run it as a service if needed. With this setup, your models can be traceable, reproducible, optimized, and easily managed without complicated operational overhead.
May 17, 2025
669 words in the original blog post.
The Lambda Cloud Metrics Dashboard is a real-time monitoring tool that provides insights into cloud GPU workloads, helping users identify performance bottlenecks and optimize resource utilization. It offers seamless, automatic, and insightful visibility into infrastructure without requiring custom scripts or heavyweight monitoring plugins. The dashboard brings clarity to the training loop by providing real-time monitoring, proactive issue detection, and optimized resource utilization. Users can deploy the Guest Agent in minutes and start making every GPU cycle count, with data security and privacy taken seriously through strict authorization protocols and publicly available source code.
May 15, 2025
623 words in the original blog post.
The Lambda company has announced a major upgrade to its Vector lineup of products with the integration of NVIDIA Blackwell architecture. This update brings significant improvements in power and performance, making it suitable for cutting-edge AI applications such as training models, inference, and multimodal research. The new Vector One desktop now features the NVIDIA GeForce RTX 5090 GPU, while the Vector desktop has been upgraded to dual NVIDIA GeForce RTX 5090 GPUs with a liquid-cooled system. The Vector Pro product line has also been updated with professional-grade NVIDIA RTX PRO Blackwell GPUs, offering up to 96 GB of GDDR7 ECC memory and MIG support. These upgrades are designed to meet the needs of researchers and ML engineers who require high-performance systems without compromising on power consumption or noise levels.
May 15, 2025
475 words in the original blog post.
The Filesystem S3 Adapter is a new feature in AWS Lambda that allows users to transfer data into or out of Lambda storage without provisioning a Virtual Machine (VM), reducing time and cost. This adapter enables a subset of the S3 API, including GetObject, PutObject, DeleteObject, and list, allowing users to use familiar S3 APIs directly with their Lambda Filesystem. The feature is designed for AI/ML practitioners and researchers using Lambda Cloud, particularly those using 1-Click Clusters, and offers benefits such as lowest-cost inference, drop-in open-source API replacement, no MLOps required, access to latest open-source models, and custom performance tuning.
May 07, 2025
572 words in the original blog post.
We recently dropped a new Security page, consolidating our existing information in one spot; but we didn't stop there: we partnered with Safebase by DRATA to launch a fully-fledged Customer Trust Portal (trust.lambda.ai) in record time. This portal is our transparency power move, offering customers and prospects access to security docs, certifications, and related materials such as SOC 2 Type II reports, pentest reports, policy documents, and more. The portal has been organized by themes like Data Security, Data Privacy, and Access Control, providing an overview of our posture plus related docs and links. Our Customer Trust Portal is a significant step forward in our commitment to transparency and trust, showcasing how we're building trust through transparency with security as our top priority.
May 02, 2025
213 words in the original blog post.
Managed Slurm on Lambda is a fully supported Slurm offering purpose-built for fast and seamless deployment on One-Click Clusters. It optimizes cluster utilization for AI/ML workloads, pre-validated on Lambda's 1 Click Cluster, and available exclusively on Lambda's 1 Click Cluster. The offering includes core Slurm capabilities such as latest Lambda-tuned Slurm config, LDAP-backed user/group management, cgroups-based resource policies, container support, Slurm roles, high availability, and pre-installed ML software modules. It also offers managed-only extras like automated Slurm patches & security updates, job history tracking, schedMD partnership for escalated issue resolution, proactive health monitoring, node-failure detection, alerting, and root-cause analysis. The offering is available in both Managed and Unmanaged flavors, with the former providing full HPC support SLAs + SchedMD backup, while the latter offers general infra support only.
May 01, 2025
636 words in the original blog post.