Home / Companies / Fireworks AI / Blog / July 2026

July 2026 Summaries

10 posts from Fireworks AI

Filter
Month: Year:
Post Summaries Back to Blog
Kimi K3 is an advanced open-source AI model hosted on Fireworks, designed to deliver high performance at a lower cost compared to closed models like Opus 5 and GPT 5.5. It excels in software development, cybersecurity, and other demanding tasks, with a zero data retention policy ensuring data privacy. The model is US-hosted and optimized for regulated industries, offering serverless deployment that eliminates the need for extensive infrastructure, allowing users to pay per token used. Kimi K3's architecture, including the Kimi Delta Attention mechanism, enhances its efficiency and performance, making it suitable for large-scale, multi-step tasks. It is praised for its cost efficiency, flexibility, and ability to handle complex workloads, setting a new standard for open models by matching or surpassing the quality of closed models. Fireworks also provides tools for fine-tuning and training Kimi K3, enabling users to customize the model to their specific needs without extensive setup, further enhancing its appeal as a cost-effective and adaptable solution for various industries.
Jul 27, 2026 1,676 words in the original blog post.
Fireworks Nexus offers a cost-effective solution for engineering teams by integrating open-weight AI models like Kimi-K3 and GLM-5.2 into existing workflows, enabling significant savings on AI expenditures without altering current engineering practices. As open models become reliable for most tasks, Fireworks Nexus addresses the operational complexity of managing these models at scale, providing centralized control over AI usage with features such as enterprise controls, cost observability, and intelligent traffic management. This platform facilitates seamless integration through FireConnect, allowing teams to maintain workflow continuity and optimize infrastructure, resulting in reduced costs and enhanced efficiency. Proven through evaluations by teams like Notion and Doximity, Fireworks Nexus demonstrates a reduction in per-merged PR costs and maintains quality by dynamically routing tasks based on complexity. With its enterprise-grade performance and flexible model management, Fireworks Nexus allows organizations to manage AI spend effectively while choosing the best models for specific tasks, ultimately offering a competitive edge in AI-powered engineering solutions.
Jul 26, 2026 867 words in the original blog post.
Kimi K3, a massive mixture-of-experts model reaching approximately 2.8 trillion parameters, has become accessible for Multi-LoRA serving and training through Fireworks Serverless Training, enabling users to fine-tune the model for specific tasks without the need for extensive infrastructure or high costs. Utilizing LoRA (Low-Rank Adaptation), which involves training a small, efficient adapter of weights rather than the entire model, users can achieve significant behavioral tuning while maintaining cost efficiency, as these adapters are easy to train and store. Fireworks offers an infrastructure that allows researchers to execute fine-tuning without the need for dedicated GPU clusters, focusing primarily on reinforcement learning tasks where models are trained to accomplish specific objectives. Two exemplar tasks, Countdown and Frozen Lake, demonstrate the model's ability to learn new behaviors and problem-solving strategies efficiently using small adapters. The training process emphasizes the importance of defining clear reward structures, as the choice of rewards significantly impacts the learning curve, illustrating that the skill lies in designing rewards that enable the model to effectively learn and achieve the desired outcomes.
Jul 26, 2026 2,014 words in the original blog post.
In a detailed exploration of the advantages of model routing over singular model use, the study evaluates Kimi K3, an open model, against Fable 5, a closed model, for their performance across approximately 1,030 tasks. The analysis reveals that routing tasks between the two models achieves 93% accuracy and is up to 50 times more cost-effective than using Fable alone, particularly for long agentic operations. The study utilized oracle routing to determine that K3 was selected for 72-96% of tasks, indicating that it is often the more cost-effective choice. While K3 and Fable display near-identical overall performance, they excel in different areas, with K3 leading in symbolic math and security tasks, and Fable in multi-language coding tasks. The research emphasizes the cost advantage of K3 due to token pricing and prompt caching, suggesting that routing tasks to the most suitable model leads to superior results both in quality and cost. This approach challenges the traditional reliance on single models, proposing that a mixture of models tailored to specific workloads yields the best outcomes, with K3 serving as a cost-efficient baseline and Fable as a specialized option.
Jul 21, 2026 1,226 words in the original blog post.
Heidi Health has partnered with Fireworks to enhance its ambient AI scribe technology, which transcribes clinician-patient encounters into professional clinical notes, significantly improving productivity by saving clinicians up to two hours per day. The collaboration aims to achieve superior performance compared to proprietary frontier models by transitioning from closed to open models, resulting in better control, performance, and cost savings. Fireworks' expertise in model fine-tuning and inference optimization has led to breakthroughs in supervised and reinforcement fine-tuning, allowing Heidi's model to outperform existing tiers in internal evaluations. The success of this partnership hinges on maintaining high data quality through aggressive filtering and utilizing larger batch sizes to stabilize training, facilitated by Fireworks AI's support for gradient accumulation. This approach demonstrates that open models can rival frontier models when subjected to rigorous Direct Preference Optimization (DPO), underscoring the importance of synthetic data and pre-evaluation strategies in enhancing AI model performance.
Jul 20, 2026 721 words in the original blog post.
Fireworks has announced a significant milestone with a $1.505 billion Series D funding round, valuing the company at $17.5 billion, supported by prominent investors such as Atreides Management, Index Ventures, and TCV. This achievement coincides with Fireworks surpassing a $1 billion annualized revenue run rate, driven by its focus on transforming general-purpose models into specialized intelligence tailored for individual companies. These models leverage proprietary data to provide competitive advantages in AI by delivering high-performance, cost-effective solutions. The company is building a future where businesses can own and optimize their intelligence, moving away from renting general intelligence. This funding will support the expansion of Fireworks' compute infrastructure and engineering team to further their mission of enabling companies to create specialized AI across various industries, as exemplified by Cursor's coding models and Harvey's legal AI.
Jul 15, 2026 271 words in the original blog post.
In the context of long-context inference, attention mechanisms are identified as the primary drivers of computational and memory costs, with sparse attention techniques like those used in MiniMax M3 being particularly effective in reducing these costs. However, implementing sparse attention efficiently is complex due to data-dependent selection and irregular memory access patterns. The Fireworks AI Performance team developed a Blackwell (SM100) kernel for M3 sparse attention, leveraging a KV-stationary execution path to mitigate these challenges. This approach involves loading each selected KV block once and attending to every query that selects it, focusing optimization on minimizing memory traffic and improving load balancing. As a result, their implementation achieves significant performance improvements, including a throughput of approximately 980 TFLOP/s at 4.1 TB/s HBM bandwidth, which represents a 1.9–2.4× speedup over a query-stationary baseline and a 1.6× improvement over MiniMax’s open-source MSA kernel. Additionally, the full module performance gains range from 1.18–1.43× over the baseline and 1.32–1.41× over open-source MSA, attributed to optimizations in memory traffic and execution scheduling.
Jul 10, 2026 2,578 words in the original blog post.
LangChain has optimized its Deep Agents harness for the NVIDIA Nemotron 3 Ultra, achieving top-tier performance among open models at a significantly reduced cost compared to closed alternatives. This enhancement, available through the Fireworks platform, allows businesses to post-train the model into specialized versions tailored to their specific workflows, ensuring competitive advantages remain proprietary. The cost-effectiveness of the model, which runs on advanced NVIDIA AI infrastructure and Fireworks' custom FireAttention kernels, is measured by cost per completed task, rather than per response, as it efficiently handles complex, multi-turn tasks. The open-stack approach, supported by NVIDIA's open model and runtime, along with Fireworks' training and inference loop, enables businesses to continually improve model performance using their data, thereby compounding their competitive edge. Since the announcement of Nemotron support, enterprises have been quick to adopt this solution for building agents in various domains, attracted by its promise of enhanced performance and cost efficiency without reliance on external proprietary APIs.
Jul 09, 2026 714 words in the original blog post.
AI teams are experiencing a pivotal moment as open-weight models become capable enough for real production workloads, granting more flexibility in building and scaling AI applications. Gumloop seized this opportunity to transition its AI agents to open-weight models, optimizing its agent harness and partnering with Fireworks AI for inference. This shift led to a 7x increase in agent chats on open-weight models within three weeks and up to 72% cost savings without compromising user experience. The company initially faced challenges with transitioning its company-wide assistant to open-weight models but succeeded in maintaining performance consistency when moving from Opus 4.8 to GLM-5.2, facilitated by Fireworks AI's reliable infrastructure. This transition underscores a broader trend in AI development, where open-weight models enhance capability, cost-effectiveness, and control, driving more companies to tailor AI agents to their specific workflows and data.
Jul 09, 2026 790 words in the original blog post.
In a recent project, an engineer successfully implemented a complex "reclaim" capability for a GPU scheduler, typically estimated as a month-long task, in just four days using the GLM 5.2 Fast model through FireConnect on Claude Code, at an inference cost of $218. This rapid development was achieved by leveraging the model's ability to quickly iterate through design, planning, and implementation phases, producing 3,000 lines of code with all unit and integration tests passing. The engineer highlighted the model's real-time feedback, which allowed for efficient problem-solving and decision-making without the typical delays associated with AI speeds and monthly token limits, thus enabling a more productive workflow. This experience underscored the transformative potential of fast open models in enhancing senior engineers' efficiency by reducing context-switching and improving focus, ultimately delivering quality results at a fraction of the usual time and cost.
Jul 07, 2026 1,334 words in the original blog post.