Home / Companies / Lambda / Blog / November 2023

November 2023 Summaries

3 posts from Lambda

Filter
Month: Year:
Post Summaries Back to Blog
The NVIDIA Transformer Engine is a cutting-edge library that accelerates transformer model performance on NVIDIA GPUs during training and inference phases. The engine leverages the capabilities of 8-bit floating point (FP8) precision on the latest NVIDIA Hopper and Ada Lovelace architecture GPUs, significantly accelerating performance while reducing memory consumption. The Transformer Engine's FP8 capabilities on the NVIDIA H100 Tensor Core GPU result in a 60% boost in performance compared to traditional FP16 operations, and achieve 3x the speed of the A100 GPU when using BF16 precision. This allows for the use of larger models that were previously constrained by memory limitations, unlocking significant performance gains and memory savings.
Nov 21, 2023 532 words in the original blog post.
Lambda Cloud Clusters are now available with the NVIDIA GH200 Grace Hopper Superchip, starting at $5.99/hr, offering dedicated GPU clusters optimized for large-scale LLM training, allowing customers to train models across thousands of GPUs with no delays or bottlenecks. The GH200 delivers up to 10X higher performance for applications running terabytes of data, and its coherent memory architecture enables unmatched efficiency and price for its memory footprint. Lambda's Cloud Cluster compute fabric leverages non-blocking NVIDIA Quantum-2 400 Gb/s InfiniBand networking, providing high throughput, low latency, and support for NVIDIA GPUDirect RDMA across the entire cluster.
Nov 13, 2023 454 words in the original blog post.
Lambda is expanding its cloud offerings to include access to NVIDIA H200 Tensor Core GPUs through Lambda Cloud Clusters, a dedicated GPU cluster service designed for machine learning teams. This collaboration enables customers to utilize the high-performance GPUs, networking, and storage required for large-scale distributed training. The NVIDIA H200 GPU offers nearly double the memory capacity of its predecessor, providing an optimal level of HBM3e memory that delivers the highest-performance model parallelism for large language models and generative AI. Additionally, the GPU's unmatched memory bandwidth of 4.8TB/s is critical for handling growing data sets and model sizes. With this expansion, Lambda customers can now access the fastest and most effective cloud infrastructure to power their largest and most demanding AI training projects.
Nov 13, 2023 373 words in the original blog post.