Home / Companies / Fireworks AI / Blog / December 2024

December 2024 Summaries

3 posts from Fireworks AI

Filter
Month: Year:
Post Summaries Back to Blog
Fireworks AI has introduced new capabilities to the DeepSeek V3 model, which already excels in reasoning and coding tasks, by adding vision capabilities through a feature called Document Inlining. The year 2024 marked significant advancements in large multimodal models, with DeepSeek V3 outperforming competitors like GPT4-o in benchmarks, particularly in coding, and achieving high scores in tests such as MMLU and BBH. Despite its existing strengths, DeepSeek V3 lacked vision capabilities, which can now be integrated using Fireworks AI's Document Inlining, allowing users to enable vision features with ease. Fireworks AI, an enterprise-scale LLM inference engine, supports the development of low-latency, high-performance generative AI applications by offering features like prompt caching and speculative API, ensuring high throughput and low total cost of ownership. Additionally, the Fireworks platform facilitates rapid deployment of open-source LLMs and supports a community of developers in building AI applications from prototype to production.
Dec 18, 2024 471 words in the original blog post.
Fireworks has launched the beta release of its speech-to-text APIs, utilizing Whisper v3-large models, which offer significant speed and cost improvements in audio transcription and translation. These APIs can transcribe one hour of audio in just four seconds, providing a low-latency experience crucial for engaging audio applications. The Fireworks Audio API includes features like transcription alignment, voice activity detection, and audio preprocessing, supporting use cases such as video captioning, speech model training, and podcast editing. The company offers two deployment methods: serverless and dedicated endpoints, with the latter providing greater scalability and production-readiness. Fireworks emphasizes the growing importance of multi-modal, audio-driven AI, showcasing compound AI systems that integrate audio with other modalities to create enriched user experiences. The service is currently free for two weeks, allowing users to explore its capabilities through a UI playground or code experimentation, with options for dedicated endpoints available for optimized performance.
Dec 09, 2024 1,307 words in the original blog post.
Cresta, backed by over $270 million in funding, is revolutionizing contact centers with its AI-driven platform, leveraging large language models to enhance agent productivity and customer interactions. To maintain its edge, Cresta uses Fireworks' scalable AI infrastructure, which provides low-latency, high-throughput support crucial for real-time applications. This partnership enables Cresta's Knowledge Assist tool to unify diverse information sources, delivering timely, context-relevant guidance to agents. By implementing a single mistral-based Ocean model cluster with LoRA adapters, Cresta achieves significant cost reductions while maintaining high scalability and customization for various customer needs. The collaboration not only addresses Cresta's infrastructure challenges but also positions the company for sustained growth and innovation in AI-powered customer engagement, as evidenced by the performance improvements of the Ocean-1 model over GPT-4 in specific tasks.
Dec 08, 2024 1,085 words in the original blog post.