June 2026 Summaries
13 posts from Fireworks AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Fireworks has launched GLM 5.2 Fast, a serverless deployment designed to enhance the efficiency and cost-effectiveness of coding agents by running 2-3 times faster than its Standard path without reserved GPUs. GLM 5.2 is optimized for agent loops that require reading, writing, and executing long-horizon tasks, facilitated by a 1M-token context window and high adaptive rate limits. The architecture employs a mixture-of-experts MLP stack and sparse MLA attention stack, allowing for parallelism tailored to different workloads, and uses prompt caching to maintain cost-effectiveness. The system supports structured outputs and maintains quality across tool-call validity and JSON-schema adherence, ensuring reliability even with faster generation speeds. Users can access it via a single API, with the option to prioritize reliability through a Priority service tier. Fast offers higher token throughput, and its seamless integration with existing workflows promises to deliver frontier-level quality and speed on a shared serverless infrastructure.
Jun 30, 2026
1,446 words in the original blog post.
Cursor has developed a specialized coding model, Composer 2, optimized for software engineering within Cursor's environment, using a combination of continual pre-training and large-scale reinforcement learning (RL) to enhance real developer workflows. Unlike general-purpose models, Composer 2 focuses on tasks relevant to software engineering, such as debugging, tool use, and multi-file code edits, achieving frontier-level performance with reduced costs and increased reliability. The model leverages Fireworks infrastructure to manage distributed RL across global clusters, facilitating inference without dedicated systems. This approach emphasizes the importance of environment fidelity, as RL's effectiveness hinges on how closely training environments replicate production conditions. By integrating production feedback into the training process, Cursor shifts AI development from static, general models to dynamic, environment-specific systems, where performance is driven by ongoing interaction and system-level improvements. The model's success underscores a broader trend in AI towards distributed RL systems that continuously evolve to meet the specific needs of their operating environments.
Jun 26, 2026
1,080 words in the original blog post.
Factory is revolutionizing software development with its agent-native system that utilizes Droids to automate the entire software development lifecycle, minimizing human intervention to where it's essential. The platform is being adopted by major companies like Adobe and Nvidia due to its unique principles of model independence and sovereign deployment, allowing enterprises to maintain control over their models, infrastructure, and data. Factory's collaboration with Fireworks provides day-zero access to open-weight models, ensuring cost efficiency and rapid deployment—a critical advantage in the fast-evolving AI landscape. The Factory Router further enhances cost management by matching tasks to the most cost-efficient models, demonstrating significant cost reductions and throughput improvements. This strategy has led to rapid adoption in the enterprise market, with Factory's open model choice becoming a key differentiator, enabling scalable and economically viable automation while avoiding vendor lock-in.
Jun 26, 2026
1,172 words in the original blog post.
Reinforcement learning on frontier models, like GLM 5.2, relies heavily on infrastructure that ensures numerical consistency between training and inference, a challenge historically managed only by top labs due to the complexity of achieving zero Kullback-Leibler Divergence (KLD) alignment. Fireworks now offers this infrastructure as a managed service, allowing broader access to this once-exclusive capability. The platform ensures batch invariance and zero-KLD train-serve alignment, which means the serving engine and trainer produce identical outputs, crucial for successful reinforcement learning that remains on-policy. This deterministic approach prevents the pitfalls of traditional methods like importance sampling and clipping, which often discard valuable learning signals. By maintaining bit-for-bit consistency across various components and under real production load, Fireworks delivers a robust system that improves learning efficiency and outcomes without sacrificing speed. This service democratizes access to advanced reinforcement learning tools, enabling enterprises and AI practitioners to harness state-of-the-art models with reliable numerics and reproducibility, a capability previously restricted to elite research labs.
Jun 24, 2026
1,951 words in the original blog post.
An open-source worker model combined with a closed-source advisor model has demonstrated improved outcomes at reduced costs across three benchmarks, including software engineering, terminal operations, and legal work. This setup uses an open-source model (Kimi-K2.6 or GLM-5.2) to handle tasks end-to-end while consulting a closed-source frontier model (Claude Opus 4.8) for review, resulting in higher success rates and cost efficiency. The advisor model only provides feedback during a single review step and cannot edit files, allowing the worker to maintain control over task execution. The study found that using an advisor increased task resolution rates with minimal additional costs, with GLM-5.2 plus advisor reaching parity or surpassing Opus as a worker in certain benchmarks at a fraction of the cost. The findings indicate the potential for scalable deployment of this worker-advisor configuration, highlighting its robustness and efficiency without the need for per-model or per-benchmark tuning.
Jun 24, 2026
1,512 words in the original blog post.
Starting July 1, 2026, Fireworks will transition all self-serve accounts to a prepaid billing system, where users will purchase credits upfront and have their usage deducted from this balance. This change aims to provide predictability in expenses, allowing users to decide how and when to add credits. Users have the option to either switch to the prepaid system immediately or wait for an automatic transition on the deadline. To avoid service disruptions, especially for production workloads, it is recommended to enable auto reload, which automatically purchases additional credits when the balance reaches a predetermined minimum. Contracted or enterprise customers will not be affected by this change and will continue under their existing billing agreements.
Jun 18, 2026
608 words in the original blog post.
Z.ai, formerly known as Zhipu and one of China's leading AI companies, has released GLM 5.2, its latest open-source model designed for long-horizon coding tasks, featuring a substantial 1M-token context window. This model has been validated using the Fireworks production stack and is touted as the strongest open-source model for coding, closing the gap with more closed models. GLM 5.2 aims to support developers by running multiple projects simultaneously without constant oversight, highlighting the importance of infrastructure in maintaining consistency over extended periods. Unlike routers that forward requests to external endpoints, Fireworks runs the model entirely on its infrastructure, ensuring full control and reliability. The model is available under an MIT license, allowing for commercial use and modification without regional restrictions, and it is being integrated into various platforms to facilitate easy adoption by developers and researchers. While public benchmarks from Z.ai and Fireworks provide general insights, the ultimate test of GLM 5.2's effectiveness lies in its performance on specific user workloads, emphasizing the need for individual evaluations and the potential for fine-tuning to meet specific needs.
Jun 16, 2026
1,112 words in the original blog post.
MiniMax M3, the latest model in the open-weight ecosystem, represents a significant advancement by combining strong agentic capabilities, native multimodality, and an extensive context window of over 500K tokens. It surpasses other models in intelligence, as evidenced by benchmarks and third-party evaluations, and offers a competitive price-to-capability ratio. The model's architecture, MiniMax Sparse Attention (MSA), enables efficient long-context processing, achieving substantial speed improvements over previous models and making real-time production latency viable. Developers can leverage M3's capabilities for complex, multi-turn coding collaboration, multimodal inputs, and long-document understanding, with the added flexibility of toggling reasoning modes. Available on Fireworks, M3 supports a range of demanding workloads and is priced competitively to facilitate adoption while plans are underway to expand its context window to 1M tokens.
Jun 12, 2026
1,160 words in the original blog post.
Alibaba has partnered with Fireworks to host and serve the Qwen 3.7 Plus, a multimodal agent model designed to handle complex agent loops by processing both text and images. Unlike previous iterations, Qwen 3.7 Plus is built not as a chat model but as an agent model capable of thinking, writing code, and taking actions via GUI or CLI, with the ability to preserve reasoning across turns for agentic tasks. This model, now available on Fireworks' serverless infrastructure, offers a pay-per-token pricing model and emphasizes scalability, performance, and data governance with a zero data retention policy. Notably, Qwen 3.7 Plus showcases significant improvements in speed and generalization over its predecessors, and its launch on Fireworks marks a shift from open-weight models to licensed ones, providing developers with independent inference infrastructure. The model is supported by both OpenAI-compatible and Anthropic-compatible APIs, and early-access opportunities for the upcoming Qwen 3.7 Max variant are also available.
Jun 12, 2026
1,189 words in the original blog post.
Moonshot's release of the K2.7 Code model, now supported on Fireworks, marks a significant advancement in coding models by using 30% fewer reasoning tokens than its predecessor, K2.6, while achieving higher scores on coding evaluations. The architecture maintains 1 trillion total parameters and a 256K context window, optimized for long-horizon agentic coding, with pricing consistent with Moonshot's public rates. This reduction in reasoning tokens is particularly valuable as it leads to shorter generations and faster loops, translating to a lower cost per completed task despite similar rate cards to K2.6. The model is available in three serving options on a serverless Fireworks platform: Standard for elastic pay-per-token usage, Priority for high-reliability traffic during peak congestion, and an upcoming Fast path for latency-constrained tasks. These enhancements cater to the bursty and uneven nature of agentic traffic, allowing developers to manage critical tasks without reserving resources or predicting traffic patterns.
Jun 12, 2026
821 words in the original blog post.
NVIDIA's Nemotron 3 Ultra is an open model designed to optimize long-running autonomous tasks such as coding agents and complex enterprise workflows, boasting 550B total parameters with a hybrid Transformer-Mamba MoE architecture and a 1M context. It offers 5x faster inference and up to 30% lower cost for agentic tasks compared to other open models, and is supported by Fireworks, a high-performance inference platform using the latest NVIDIA GPUs and proprietary optimizations for increased throughput. Nemotron 3 Ultra is available for both inference and post-training on Fireworks, allowing seamless transition from training to production without system handoffs, and can be customized through supervised fine-tuning and direct preference optimization. The platform provides on-demand deployments with dedicated GPUs and predictable performance, making it cost-effective and efficient for enterprises looking to enhance their AI capabilities.
Jun 04, 2026
531 words in the original blog post.
In a study examining the performance of legal AI models using Harvey's Legal Agent Benchmark (LAB), the integration of open-source models with frontier tools and Fireworks-native post-training techniques significantly improved performance and cost efficiency. The hybrid system, featuring an open-source GLM 5.1 worker and Claude Opus 4.7 as an advisor, achieved an 18/100 all-pass rate at a reduced cost of $368, outperforming Opus alone, which had a 14/100 rate at $954. Post-training on the Fireworks platform, utilizing supervised and reinforcement fine-tuning on models like Kimi K2.6, further enhanced performance, demonstrating the potential of open-source models to approach frontier-level quality while maintaining cost-effectiveness. This approach allowed for a seamless transition from research to production, with no discrepancies between training and serving models, emphasizing the competitive edge of open-source solutions in legal AI tasks.
Jun 03, 2026
2,368 words in the original blog post.
Trilogy’s AI Center of Excellence faced rising infrastructure costs and operational challenges as AI adoption grew across its portfolio companies, prompting them to standardize and optimize their AI workflows. To address these issues, they adopted Fireworks AI as the primary inference infrastructure, which allowed the transition from fragmented model experimentation to scalable, production-grade usage of open-weight models. This shift facilitated billion-token scale workloads, reduced costs, and improved operational efficiency by supporting structured evaluation and testing within a unified environment. Fireworks enabled rapid iteration across teams, facilitating the deployment of open-weight models without significant infrastructure overhead, and became integral to systems like OpenSymphony, which orchestrates high-volume, multi-agent workflows. Consequently, Fireworks empowered Trilogy to move from isolated experimentation to enterprise-scale agentic systems, significantly enhancing the efficiency and scalability of their AI operations.
Jun 01, 2026
1,206 words in the original blog post.