September 2025 Summaries
6 posts from Lambda
Filter
Month:
Year:
Post Summaries
Back to Blog
Lambda is at the forefront of a transformative shift in data center infrastructure, emphasizing the need for facilities optimized for GPUs rather than traditional servers to meet the increasing demands of AI-scale workloads. Kenneth "Ken" Patchett, VP of Data Center Infrastructure at Lambda, highlights the company's innovative approach to building AI factories that prioritize fast deployment, density, and advanced cooling systems to maximize efficiency and intelligence per watt. These next-generation facilities operate at significantly higher power densities of 130 to 240 kW per rack, as opposed to the traditional 2 to 15 kW per rack, necessitating advancements in cooling technology and infrastructure design. Lambda offers both public and private cloud solutions, allowing enterprises to lease GPUs on demand or through dedicated contracts, enabling them to focus on AI model training without the complexities of high-performance compute orchestration. Patchett argues that the bottleneck in AI advancement is not GPU scarcity but the lack of data centers equipped for AI-scale tasks, emphasizing the importance of modular design and collaboration in infrastructure development. Looking forward, Lambda aims to democratize compute power with the vision of "one GPU per person," supporting innovation in various fields such as healthcare and education while addressing sustainability and energy efficiency challenges.
Sep 30, 2025
4,292 words in the original blog post.
Lambda has appointed Heather Planishek to its Board of Directors as Audit Chair, bringing her extensive experience in scaling high-growth technology companies, notably as COO and CFO at Tines and as Chief Accounting Officer at Palantir Technologies. Her role will involve overseeing Lambda’s financial reporting, risk management, and compliance as the company continues to expand its operations in AI infrastructure, serving major AI labs, hyperscalers, and over 200,000 AI developers. Planishek, recognized for her strategic and executional acumen, previously worked at Hewlett Packard Enterprise and EY, and she joins Lambda at a pivotal time as it aims to make AI infrastructure as ubiquitous as electricity. Founded in 2012, Lambda is focused on building large-scale AI factories and providing essential infrastructure for AI development and deployment, with a mission to democratize access to artificial intelligence.
Sep 25, 2025
389 words in the original blog post.
Lambda, in collaboration with ECL and Supermicro, has launched the first hydrogen-powered NVIDIA GB300 NVL72 systems at ECL's Mountain View facility, marking a significant step in sustainable AI infrastructure. The Supermicro-built systems receive 142 kW of compute power and are cooled using direct-to-chip liquid systems, with the data center operating entirely on hydrogen fuel cells, achieving zero emissions and water usage. Lambda has expanded its presence at ECL to fully utilize the facility, emphasizing its commitment to hydrogen as a sustainable power source for AI factories. This deployment demonstrates the feasibility of off-grid, zero-emission data centers at scale, with the systems integrated into the facility in record time, reflecting a major industry milestone. The collaboration underscores the potential for hydrogen power to support AI infrastructure growth responsibly, aligning with Lambda's mission to provide scalable and sustainable superintelligence capabilities.
Sep 23, 2025
772 words in the original blog post.
JAX is becoming a preferred framework for machine learning teams seeking both research flexibility and production performance, particularly when utilizing GPUs. Unlike PyTorch, which uses an imperative style, JAX's functional approach combined with XLA compilation offers significant advantages for GPU workloads, including automatic kernel fusion and multi-accelerator scaling. This guide provides strategies for selecting the right framework, configuring GPU environments, optimizing core operations, and scaling across multiple GPUs, as well as best practices for production deployment. It contrasts JAX's capabilities with PyTorch, highlighting JAX's strengths in hardware portability, composable transformations, and memory efficiency, while noting PyTorch's mature ecosystem and intuitive style. The guide also covers essential setup steps, core optimization techniques, and scaling strategies for transitioning from single to multi-GPU deployments, alongside performance best practices like gradient checkpointing and mixed precision. By following these strategies, teams can effectively deploy JAX on GPUs, optimizing memory usage and building efficient production pipelines, with Lambda providing the necessary GPU infrastructure for enterprise machine learning workflows.
Sep 19, 2025
1,409 words in the original blog post.
Lambda and Cologix have introduced the first NVIDIA HGX B200 AI cluster in Columbus, Ohio, marking a significant step in decentralizing AI infrastructure away from traditional coastal hubs. This deployment leverages Lambda's one-click cluster technology, Supermicro's energy-efficient systems, and Cologix's carrier-dense facilities to provide scalable and flexible GPU access, which is reshaping AI adoption across various industries. Columbus is strategically chosen for its central location, rich network connectivity, and favorable economic conditions, making it an ideal site for AI workloads that demand geographic diversity without compromising performance. The collaboration aims to lower barriers to AI entry for both large and small enterprises, enabling quicker deployment and scaling of AI applications. As AI hardware evolves rapidly, Cologix and Lambda are adapting data center designs to meet future needs, incorporating both air and liquid cooling capabilities to support high-density, high-performance computing. This partnership represents a repeatable model for regional AI cluster development, poised to influence sectors like healthcare, logistics, and manufacturing by bringing AI capabilities closer to the point of need.
Sep 09, 2025
5,868 words in the original blog post.
Lambda's MLPerf Inference v5.1 results demonstrate significant performance improvements, with up to 15.4% gains over prior benchmarks, showcasing the capability of NVIDIA HGX B200-powered 1-Click Clusters to enhance enterprise inference workloads. The results highlight the performance of models like Llama 2 70B, Llama 3.1 405B, and Stable Diffusion XL across different scenarios, with the Llama 3.1 405B model achieving notable server-side gains. These benchmarks were achieved using NVIDIA's latest technologies, including TensorRT 10.11 and CUDA 12.9, emphasizing not only hardware advancements but also software optimizations. The tests were conducted on a consistent system configuration, focusing on maximizing throughput and minimizing latency under real-world conditions. Lambda's infrastructure, designed for enterprise AI, supports scalable GPU clusters with flexible rental terms, making it suitable for both startups and enterprises looking to validate AI use cases or scale their operations.
Sep 09, 2025
782 words in the original blog post.