Home / Companies / Lambda / Blog / December 2023

December 2023 Summaries

3 posts from Lambda

Filter
Month: Year:
Post Summaries Back to Blog
The NVIDIA GH200 Grace Hopper Superchip, a GPU-CPU hybrid system, has been benchmarked with ZeRO-Inference technology to demonstrate its potential in handling large AI models. The results show that the combination of ZeRO-Inference and the GH200 superchip effectively handles LLMs up to 176 billion parameters, significantly improving inference throughput compared to standalone GPUs like H100 or A100 Tensor Core GPU. The high bandwidth NVLink-C2C interconnect and Address Translation Services enable seamless memory access for both CPU and GPU, facilitating efficient inference. ZeRO-Inference reduces costs of large AI models by offloading model weights to CPU memory or NVMe, making advanced AI models more accessible. The benchmark results demonstrate the benefits of leveraging the GH200's high bandwidth for CPU-offload, producing higher throughput with larger batch sizes and enabling running inference with larger models than before. Overall, this combination marks a major leap in AI inference technology, democratizing access to advanced AI models and opening new possibilities for computational efficiency and scalability.
Dec 20, 2023 434 words in the original blog post.
Persistent storage for on-demand NVIDIA H100 GPU instances is now available in all Lambda Cloud regions and instance types, including persistent storage on the 8x and 1x varieties of this GPU. This feature allows customers to persist files and data when using Lambda's compute service without incurring additional costs or quotas. The cost of storage for Lambda Cloud instances is $0.20/GB/month, with some exceptions, such as free storage in the Texas region until the end of 2023. Filesystems can be attached to instances but not across regions, and data cannot be accessed without attaching to a VM. Customers are advised to download anticipated data before shutting down their instance.
Dec 19, 2023 288 words in the original blog post.
The Lambda Vector One is a new single-GPU desktop PC designed to tackle demanding AI/ML tasks. It features a powerful NVIDIA GeForce RTX 4090 graphics card, an AMD Ryzen 9 7950X CPU, and up to 128 GB of DDR5 memory. The system is built with liquid cooling for optimal performance without noise, making it perfect for quiet workspaces. The Vector One comes pre-installed with Lambda Stack, which includes everything needed to train neural networks, and offers a one-year warranty on hardware with the option to extend to three years. With its powerful architecture and optimized configuration, the Vector One is ideal for deep learning tasks and meets budget requirements without sacrificing performance.
Dec 12, 2023 493 words in the original blog post.