June 2025 Summaries
5 posts from Lambda
Filter
Month:
Year:
Post Summaries
Back to Blog
Iambic Therapeutics has partnered with Lambda to enhance its AI-driven drug discovery efforts, utilizing Lambda's NVIDIA HGX B200 cluster to support the training of Enchant, a groundbreaking multi-modal transformer model. Enchant predicts clinical and preclinical endpoints, allowing researchers to evaluate new drug molecules and make confident predictions even with limited data. The latest version, Enchant v2, provides accurate predictions for crucial properties related to drug development, potentially improving the success rate of drug candidates in clinical trials. This collaboration is expected to accelerate breakthroughs in life sciences by optimizing drug design and clinical trial planning. Iambic's AI-driven discovery platform, which includes innovations like Enchant and NeuralPLexer, integrates physics principles to improve data efficiency and explore a wide range of chemical structures. Iambic, founded in 2020 in San Diego, is advancing a pipeline of novel medicines to address unmet patient needs, while Lambda, established in 2012, aims to become the leading AI compute platform, offering extensive GPU cloud services for developers.
Jun 26, 2025
756 words in the original blog post.
Apriel 5B is a compact yet powerful large language model (LLM) developed by ServiceNow, designed to be efficient and cost-effective for enterprise deployment. With 4.8 billion parameters and a decoder-only transformer architecture, it focuses on maximizing performance while minimizing compute costs and latency, making it suitable for high-throughput inference and fine-tuning tasks. Trained on a diverse dataset of 4.5 trillion tokens, including natural language text and programming languages, Apriel 5B excels in various enterprise applications such as IT service management automation, conversational AI, code generation, and domain-specific natural language processing. Despite its smaller size compared to larger models, it achieves an impressive inference throughput of approximately 1,250 tokens per second. Apriel 5B supports mixed precision quantization, pipeline and tensor parallelism, and is ready for production use through ONNX Runtime and Triton Inference Server deployments. Its development was supported by Lambda's infrastructure, facilitating experimentation and iteration without the constraints of hardware limitations or disruptions.
Jun 20, 2025
1,186 words in the original blog post.
Lambda and dstack provide a streamlined alternative to Kubernetes and Slurm for teams running on Lambda, with native support for training workloads, development environments, and persistent services. By using Lambda's 1-Click Clusters and dstack's orchestration, developers can focus on building rather than setting up their machine learning infrastructure. The RAGEN framework is used for training large language models as reasoning agents in complex, multi-turn environments, introducing fine-grained reward signals to improve agent reliability and performance. With dstack, users can launch a Ray cluster with the RAGEN environment, submit Ray tasks from their local machine, and recover training in case of a failure or cluster restart.
Jun 05, 2025
1,011 words in the original blog post.
DeepSeek-R1-0528, an open-source model, has been released on Lambda's Inference API, challenging the dominance of OpenAI's o3 and Google's Gemini 2.5 Pro in complex tasks. The latest release builds upon the deepseek_v3 architecture, employing FP8 quantization to enhance its capabilities. At its core lies a robust architecture based on the DeepSeek-V3 backbone, utilizing a mixture-of-experts (MoE) model with multi-headed latent attention (MLA) and multi-token prediction (MTP). This approach enables efficient handling of complex reasoning tasks and allows the model to learn and improve through trial and error. R1-0528 demonstrates notable improvements over its predecessor across various benchmarks, achieving an impressive 87.5% accuracy in the AIME 2025 benchmark and scoring 73.3% on LiveCodeBench. The model now supports JSON output and function calling, enhancing its utility in various applications. It has significantly suppressed hallucination issues present in the legacy R1 version, leading to reliable and consistent outputs. With its sophisticated architecture, increased token utilization, and reduced dependence on supervised datasets, DeepSeek-R1-0528 is positioned to lead the next wave of AI advancements.
Jun 04, 2025
832 words in the original blog post.
Cologix and Lambda have collaborated to deploy NVIDIA HGX B200-accelerated AI clusters at Cologix's COL4 ScalelogixSM data center in Columbus, Ohio. This is the first deployment of its kind in the region, delivering enterprise-grade AI compute with simplified access for businesses across the Midwest. The collaboration brings fast and cost-effective ways to support large model training, fine-tuning, and inference workloads, enabling regional enterprises to spin up high-performance compute infrastructure in seconds without infrastructure management required. Built on Supermicro's AI-optimized hardware and Cologix's carrier-dense environment, this deployment increases the availability of Lambda's popular 1-Click Clusters, which are purpose-built for model training and inference at scale. The launch reflects a broader trend of moving high-performance compute closer to where data is generated and used, enabling enterprises in healthcare, finance, logistics, retail, and manufacturing to accelerate their AI roadmaps without the heavy lift of managing infrastructure.
Jun 03, 2025
980 words in the original blog post.