October 2025 Summaries
16 posts from Fireworks AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Genspark, an innovator in AI-powered applications, has successfully enhanced its Deep Research Agent by collaborating with Fireworks AI, achieving significant improvements in quality and tool calls while reducing costs. Through the use of Fireworks' Reinforcement Fine Tuning (RFT), Genspark fine-tuned large open-source models, resulting in a 12% improvement in quality and a 33% increase in tool calls compared to state-of-the-art closed-source models, leading to a 50% reduction in costs. This collaboration allowed Genspark to overcome challenges associated with closed-source models, such as limited customization and scalability, by adopting a mixture of agents tailored to specific workloads. The partnership facilitated a seamless transition from training to deployment using advanced infrastructure, including NVIDIA hardware, and allowed Genspark to offer superior user experiences by delivering comprehensive, multi-source research reports. The support from Fireworks also enabled Genspark to rapidly transition from proof of concept to production within four weeks, demonstrating the collaborative effort's success in optimizing AI capabilities and enhancing user engagement.
Oct 31, 2025
1,063 words in the original blog post.
Fireworks AI has successfully secured $250 million in Series C funding, elevating its valuation to $4 billion, with involvement from major investors such as Lightspeed Venture Partners, Index Ventures, and Sequoia Capital. The company, founded by former PyTorch engineers, aims to empower enterprises to develop, own, and customize their AI infrastructure, moving away from dependency on a few major foundation model providers. Fireworks AI's platform supports over 10,000 companies, including high-profile clients like Samsung, Uber, and Shopify, by providing access to a wide range of open-source models and facilitating the development of production-grade AI applications. This funding will enable Fireworks AI to expand its research in tuning and inference alignment, enhance its AI creation toolchain, and scale its global compute infrastructure, reinforcing its position as a leader in AI inference technology. The company's approach focuses on "one-size-fits-one" AI, allowing enterprises to utilize unique data to fine-tune models for specific use cases, thereby continuously improving their AI applications through a feedback loop of product-model co-design.
Oct 28, 2025
788 words in the original blog post.
Fireworks AI has raised $250 million in Series C funding to advance enterprise artificial intelligence, particularly through the integration of NVIDIA's Nemotron Nano2 VL model. This cutting-edge vision language model (VLM) is designed to enhance document intelligence and video understanding applications by combining a large language model (LLM) with a vision encoder, allowing AI systems to extract and interpret information across multiple modalities such as text, images, tables, and video. Nemotron Nano2 VL's capabilities include high accuracy in tasks such as optical character recognition, chart reasoning, and video comprehension, making it a versatile tool for automating workflows in industries like finance, healthcare, and government. The model's efficient Mamba-Transformer architecture and open-source nature provide flexibility and cost-effective scalability, as demonstrated by its success in automating invoice processing with over 90% accuracy. Fireworks AI offers resources to help users deploy this technology and unlock insights from complex documents and multimedia content, thereby reducing operational costs and enhancing productivity.
Oct 27, 2025
793 words in the original blog post.
Fireworks has introduced Deployment Shapes to streamline the configuration of serving setups for developers using large language models (LLMs). These pre-configured templates are designed to optimize deployments for latency, throughput, or cost, balancing the other factors to suit different use cases. Users can start with serverless deployments, which are easy to use but may not be optimal for high-volume needs, or opt for on-demand deployments that offer single-tenant, customizable configurations. Fireworks' advanced techniques, such as speculative decoding and caching, enhance inference speed and efficiency, while ongoing improvements in GPU kernels and configurations ensure cutting-edge performance. Deployment Shapes are now available via both the Fireworks website and CLI, and the company offers additional customization support for enterprise customers seeking further optimization.
Oct 23, 2025
828 words in the original blog post.
Fireworks and AMD have formed a multi-year strategic partnership to enhance the capabilities of AMD Instinct™ GPUs, aiming to accelerate their adoption among AI-native companies, developers, and enterprises. This collaboration combines AMD's expertise in high-performance computing with Fireworks' advanced AI stack to create scalable, cost-efficient AI systems that improve inference speed and quality. By optimizing Fireworks' software stack for AMD's MI325X and MI355X accelerators, the partnership aims to lower the total cost of ownership, enhance throughput and latency, and facilitate rapid deployment at scale. The agreement provides Fireworks with access to AMD's cutting-edge accelerators through preferred cloud service providers, ensuring that advancements in AMD's hardware lead to better performance and efficiency for Fireworks' customers. This partnership is set to drive the next generation of AI infrastructure, emphasizing faster, more efficient, and open-source AI innovation.
Oct 20, 2025
341 words in the original blog post.
The blog post discusses the use of Fireworks Eval Protocol and Ollama to facilitate the selection and deployment of AI models by allowing teams to replace OpenAI models with local open-source alternatives without altering existing application logic. By maintaining a consistent evaluation harness, teams can seamlessly swap model backends using an OpenAI-compatible API provided by Ollama, enabling robust, evidence-based model comparisons. Two examples are provided: one involves evaluating an agent on the Chinook dataset using PydanticAI, and the other assesses Langfuse traces, demonstrating how local models like qwen3:8b can outperform some remote models in specific tasks. The approach supports minimal code changes, preserving evaluation and logging processes while enabling rapid and secure validation of alternative models.
Oct 15, 2025
852 words in the original blog post.
Fireworks has announced an upgrade to its platform for Retrieval-Augmented Generation (RAG) workloads, introducing the Qwen3 8B Embeddings and Reranking models to a serverless environment, along with two new API endpoints for seamless access. These advancements aim to simplify the construction of scalable RAG applications by supporting open models for each step of the process—embedding, indexing, retrieving, reranking, and synthesizing—on a unified platform, eliminating the need for multiple providers. The platform's enhancements include top-tier performance, global scalability, a consistent developer experience, unified billing, and an expanded model library, supporting various BERT-based embeddings models. Additionally, Fireworks encourages user engagement for future developments, inviting feedback to help shape its roadmap and improve features further.
Oct 09, 2025
870 words in the original blog post.
Fireworks has announced the release of Streaming Transcription V2 and Streaming Speaker Diarization, enhancing real-time speech-to-text capabilities and speaker identification in audio streams. Streaming Transcription V2 offers a faster, lower-latency API, improving on its predecessor with up to 25% reduced latency and better accuracy in noisy environments, all at a cost-effective price. This upgrade is crucial for applications like live captioning and customer support automation, where immediate transcription is necessary. Meanwhile, the new Streaming Speaker Diarization, now in closed beta, provides real-time speaker identification, maintaining consistent speaker IDs and offering flexible integration with transcription results, which is useful for call center analytics and live broadcasts. Both tools are designed to support high-volume concurrent streams, offering scalable and reliable solutions for interactive voice agents and other AI-driven audio applications.
Oct 06, 2025
789 words in the original blog post.
Fireworks has introduced Document Inlining, a system designed to address the challenges of processing multimedia content by converting various digital asset formats, such as PDFs and images, into text that Large Language Models (LLMs) can easily process. This solution aims to overcome the limitations of Vision Language Models (VLMs) that often struggle with non-textual data, resulting in reduced reasoning capabilities and increased costs. Document Inlining automates the transformation of documents into a structured text format, enabling LLMs to process and reason with this data effectively. By using a specialized parsing service, it handles complex document structures like tables and charts, enhancing the quality of results and improving processing speed through parallel transcription. Fireworks' approach allows for flexible input types, improved quality through specialized components, and ultra-simple usage compatible with the OpenAI API. The system has been shown to deliver superior performance compared to other models and promises to extend its capabilities to include audio inlining and long document searches in the future.
Oct 06, 2025
1,685 words in the original blog post.
Fireworks has developed a platform for creating customizable, real-time voice agents that integrates automatic speech recognition (ASR), text-to-speech (TTS), and large language models (LLM) into a single, efficient solution. This approach addresses the challenges of traditional cascaded systems, such as latency, cost, and complexity, by co-locating components and using optimization techniques to achieve sub-500 millisecond response times. Fireworks' platform offers accurate ASR, capable of handling accents and background noise, and crisp TTS that allows for specific pronunciation and voice customization. The platform also supports advanced LLM capabilities for following complex instructions and integrating tool calls, offering a fully customizable end-to-end experience. Currently in beta, Fireworks invites users to explore its capabilities and provides a free, limited-access endpoint for demonstrations, encouraging potential design partners to collaborate on optimizing voice agent stacks.
Oct 06, 2025
889 words in the original blog post.
Fireworks has enhanced its Whisper-based speech transcription service by introducing new features, including speaker diarization and a Batch API, to address customer demand for more sophisticated audio processing capabilities. The diarization feature identifies individual speakers in audio recordings, providing valuable insights for applications like meeting transcription and phone call analytics, while maintaining high scalability and accuracy. The Batch API allows users to process large volumes of audio files cost-effectively, offering a 40% reduction in price compared to typical APIs and returning results within 24 hours, making it ideal for use cases that do not require immediate responses. These improvements, combined with Fireworks’ existing audio services, enable the development of advanced AI applications for contact-center analytics, media indexing, and more, by integrating speech, text, and other modalities into a cohesive AI pipeline. Fireworks aims to provide a robust platform for building, customizing, and scaling AI systems with flexible deployment options and ongoing support through community channels like Discord and Twitter.
Oct 06, 2025
1,362 words in the original blog post.
Fine-tuning large language models (LLMs) is crucial for adapting general-purpose models to meet enterprise-specific requirements such as precision, compliance, and reliable outputs. Unlike pre-training, which equips models with broad language understanding, fine-tuning updates a pre-trained model's weights using specialized datasets to align it with niche domains such as healthcare, finance, or legal sectors. There are several approaches to fine-tuning, including full fine-tuning and parameter-efficient methods like LoRA, which balance computational cost and effectiveness. Fine-tuned models significantly improve accuracy, reduce error rates, and provide consistent structured outputs, making them indispensable in environments that require strict adherence to domain-specific terminology or regulatory standards. Fireworks AI offers a robust platform for fine-tuning, providing tools for efficient training, deployment, and continuous evaluation, thus helping organizations transition from experimental models to scalable, enterprise-grade AI systems.
Oct 06, 2025
1,976 words in the original blog post.
Fireworks has introduced a new streaming speech-to-text API designed for real-time applications such as voice agents and live captioning, featuring an impressive 300ms end-to-end latency for 16kHz mono PCM audio and accuracy within 3% WER of Whisper v3-large. The API is cost-efficient, priced at $0.0032 per audio minute, making it significantly cheaper than competitors. The service is tailored for immediate transcription needs in scenarios like call centers and live broadcasts, providing incremental text segments via a WebSocket connection that streams audio chunks of 50-500ms intervals. Fireworks' custom audio serving stack, built over years with Pytorch, employs optimizations like voice activity detection to manage sparse speech audio efficiently. In addition to speed and accuracy, the service offers production-readiness, supporting companies like Cursor, Uber, and Doordash, and providing serverless customers with a quota of 50 concurrent streams. Fireworks also enables broader compound AI systems by integrating speech with text, image, and specialized models, offering flexibility and adaptability for diverse use cases. Users can begin with the serverless streaming endpoint through code or a UI playground, facilitating ease of use and integration into existing workflows.
Oct 06, 2025
1,119 words in the original blog post.
A global quick-service restaurant (QSR) chain is transforming its drive-thru operations by implementing Fireworks' real-time voice intelligence platform to enhance customer interactions and operational efficiency. The initiative aims to digitize order-taking entirely, enabling faster and more personalized service. Fireworks provides a unified platform delivering sub-500ms transcription latency, significantly reducing costs and improving accuracy even in noisy environments. This solution has demonstrated success with a rapid pilot-to-production rollout, showing a 3-5 times return on investment in initial stores and plans to scale across over 6,000 locations. This strategic partnership, validated through technical benchmarking against competitors like OpenAI and Whisper, positions the brand to leverage voice AI for smarter upselling and enhanced customer satisfaction. Future plans include expanding Fireworks' capabilities to voice agents and analytics, supporting the QSR's vision for a fully digital drive-thru experience.
Oct 06, 2025
1,019 words in the original blog post.
Fireworks AI has integrated with AWS AgentCore to enable developers to efficiently deploy AI agents with optimized inference on AWS's secure, serverless infrastructure. AWS AgentCore provides serverless scaling, built-in security, and flexibility to use any open-source framework, simplifying the deployment of dynamic AI agents at scale without the need for container management. Fireworks AI enhances this by offering an advanced inference engine with features like adaptive caching and speculative decoding to ensure fast, sub-second latency in multi-turn agent interactions. This integration allows developers to build and deploy AI agents locally and globally with enterprise-grade security, leveraging AWS's existing infrastructure and compliance protocols. An example provided is a code generation agent built using AgentCore, which demonstrates the seamless deployment and operation of AI agents from local development to global deployment via AWS CodeBuild. The collaboration promises continuous expansion with more comprehensive platform integrations and examples to support complex AI agent workflows.
Oct 02, 2025
401 words in the original blog post.
Fireworks for Startups is a newly launched program designed to assist startups in developing AI products by providing a comprehensive suite of tools, expert support, and community engagement to expedite innovation and scaling. The program addresses common challenges faced by startups, such as managing infrastructure and costs, by offering access to advanced AI models, world-class engineering support, and a robust platform for efficient and reliable scaling. Participants benefit from direct interactions with AI experts, access to a dedicated startup community, and resources like libraries and guides. The program also offers flexible pricing models and infrastructure that supports rapid prototyping and production-level scaling, enabling startups to focus on creating differentiated and market-ready AI solutions without the burden of technical complexities. Additionally, Fireworks facilitates brand visibility and collaboration through joint marketing efforts and case studies, further supporting startups in their growth and product development journey.
Oct 01, 2025
473 words in the original blog post.