Home / Companies / Fireworks AI / Blog / March 2025

March 2025 Summaries

3 posts from Fireworks AI

Filter
Month: Year:
Post Summaries Back to Blog
Fireworks AI is offering a comprehensive developer platform designed to empower developers with advanced AI tools, focusing on speed, efficiency, and cost-effectiveness. Central to their offerings is the DeepSeek R1 model, which has been optimized for high-speed performance and cost efficiency through innovations like FireAttention and a specialized inference engine. New deployment options are available on Hopper GPUs, with plans to include Blackwell GPUs, enhancing both speed and throughput. The platform provides secure hosting, model quality customization, and agentic development capabilities, including multi-modal workflows and seamless tool integrations. Fireworks AI is committed to delivering a transparent, steerable, and privacy-focused toolchain that supports real-time, low-latency interactive experiences, catering to developers' varied needs in AI product development.
Mar 18, 2025 386 words in the original blog post.
Fireworks AI has announced support for NVIDIA NIM microservices, part of the NVIDIA AI Enterprise software platform, enhancing its capability to deploy AI models with increased speed, customization, and cost efficiency. The integration allows enterprises to leverage NVIDIA's advanced AI models, such as DeepSeek and Llama, on Fireworks' platform, enabling faster AI inference processing and seamless model customization. This development supports a wide array of AI applications, including text generation, image and vision, embeddings and search, and audio processing, all optimized for multi-model workflows. By integrating Fireworks AI's Compound AI framework with NVIDIA's specialized microservices, the collaboration promises to accelerate complex tasks like drug discovery through a compound system architecture comprising modules for protein structure analysis, molecular design, similarity-based screening, and binding validation. This unified ecosystem of models and architectures is designed to offer scalable, efficient solutions for various industry applications, making advanced AI more accessible and actionable for businesses.
Mar 18, 2025 851 words in the original blog post.
Fireworks has introduced Quantization Aware Fine Tuning (QAT) for its DeepSeek R1 and V3 models, aiming to optimize these state-of-the-art open models for quality, latency, and cost through their FireOptimizer adaptation engine. The challenges of fine-tuning these models include dealing with accuracy drops from varying training and serving configurations, substantial GPU memory requirements due to their 671 billion parameters, and complexities with the Mixture-of-Experts structure in DeepSeek V3. QAT, which builds on LoRA and QLoRA techniques, helps achieve high accuracy with reduced memory usage by simulating inference setup through "fake quantization" of merged weights and activations. This method is shown to provide an edge over naive FP8 LoRA tuning and can be seamlessly implemented on Fireworks for models like Llama and DeepSeek V2, promising faster inference speeds without significant accuracy loss. The initiative also highlights the stability and improved alignment of evaluation metrics with inference numerics, showcasing the potential of QAT for expanding model capabilities beyond DeepSeek to other bfloat16 models.
Mar 12, 2025 890 words in the original blog post.