April 2025 Summaries
2 posts from Lambda
Filter
Month:
Year:
Post Summaries
Back to Blog
The DeepSeek V3-0324 endpoint is a new AI development platform that offers lightning-fast responses, up to 128K massive context window, and no rate limiting, all for the low price of $0.88 per 164K output. It features a 685B parameter model with a Mixture-of-Experts (MoE) design and has been trained on 14.8 trillion tokens using an auxiliary-loss-free load balancing strategy and multi-token prediction (MTP). The endpoint outperforms other models in structured reasoning and creative tasks, achieving high scores in benchmarks such as MATH-500, Massive Multitask Language Understanding(MMLU Pro), GPQA Diamond, AIME, and LiveCodeBench. With its ease of integration on the Lambda Inference API, developers can quickly get started with DeepSeek v3-0324 and unlock open-source inference without artificial limitations.
Apr 19, 2025
411 words in the original blog post.
Lambda has released its first public-facing results on NVIDIA HGX B200 and H200 platforms for MLPerf Inference v5.0, showcasing the performance of its innovative cloud infrastructure. The platform harnesses the power of NVIDIA accelerators to deliver state-of-the-art AI computing capabilities. The tests demonstrate significant throughput improvements, scalability, and reliability, with the 8-GPU node sustaining high performance and completing samples faster than previous rounds. Lambda's commitment to providing the best compute platform for AI innovation is evident in its collaboration with NVIDIA and its ongoing efforts to optimize software and scalable AI infrastructure. With its cloud solutions engineered to support ambitious AI development projects, Lambda aims to make state-of-the-art AI computation accessible to everyone, pushing the boundaries of what is possible and driving innovation in the field.
Apr 02, 2025
540 words in the original blog post.