June 2025 Summaries
8 posts from Arcee AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Arcee AI has released the open weights of five language models, including three enterprise-grade production models and two research models, to support open collaboration and democratize AI technology access. These models, previously available only through the Arcee Conductor SaaS platform, are now accessible for unrestricted deployment, allowing developers, researchers, and enterprises to customize and build upon them for various applications, such as customer service automation, technical documentation, and API orchestration. The production models, Arcee-SuperNova-v1, Caller, and Virtuoso-Large, have been tested in real-world enterprise environments and come with commercial licensing, while the research models, GLM-4-32B-Base-32K and Homunculus, offer insights into advanced training techniques and architecture innovations. This release aligns with Arcee AI's transition toward focusing on the Arcee Foundation Model (AFM) family, promoting transparency and community-driven innovation in AI development.
Jun 30, 2025
968 words in the original blog post.
MergeKit is a leading tool for Model Merging, which combines multiple pre-trained models into a single, more efficient one while preserving their original capabilities and enhancing performance. This technique, explored by Arcee in various research papers, reveals its potential not only in post-training but also during the pre-training phase, as demonstrated by the introduction of Pre-trained Model Average (PMA) that stabilizes the training process and improves model performance. In healthcare, the PatientDx framework utilizes model merging to create domain-specific models for predictive tasks without compromising patient data privacy. Additionally, the method proves effective in adapting language-specific large language models (LLMs) to enhance reasoning capabilities for low-resource languages through a strategic merging process. These studies underscore the versatility and cost-effectiveness of model merging in developing robust, domain-specific AI models across different industries and applications.
Jun 24, 2025
1,560 words in the original blog post.
Arcee has unveiled its new foundation model, AFM-4.5B, which boasts an extended context length from 4k to 64k through a rigorous and experimental training process. This model, intended for both short and long-context tasks, is a product of various innovative techniques including model merging, distillation, and context extension strategies derived from existing research. Initial evaluations reveal promising results, though these are preliminary and subject to change as further training continues. AFM-4.5B's development involved a series of experiments, leveraging methods like YaRN positional embeddings, ProLong training, and model merging to enhance performance while maintaining short-context task efficiency. The model's final form achieved notable improvements in benchmarks such as MMLU and Big Bench Hard, demonstrating a balance between maintaining short-context accuracy and extending long-context capabilities. The process highlights the effectiveness of linear averaging and distillation in refining model performance, suggesting that these methods can scale to larger models as well, as evidenced by successful trials on the GLM-32B base model.
Jun 23, 2025
2,752 words in the original blog post.
Arcee AI has launched AFM-4.5B, its first foundation model, designed to meet the demands of modern enterprises by offering high performance, compliance, and affordability at a scale previously unavailable. In response to customer challenges with existing AI models, AFM-4.5B was created to bridge performance gaps and regulatory issues, employing rigorous data curation and post-training strategies to enhance reliability and flexibility. The model was trained using 6.58 trillion tokens of curated data through a partnership with DatologyAI and leveraged Amazon SageMaker's infrastructure for rapid experimentation. AFM-4.5B is optimized for cost-effective inference and supports deployment from cloud to edge devices, making it suitable for a wide range of enterprise applications. The model's post-training pipeline, incorporating advanced techniques like reinforcement learning and alignment methods, ensures high accuracy and adaptability across various tasks. The preview version is available for testing, and Arcee AI plans to release the final model on Hugging Face, promoting transparency and community involvement in AI development.
Jun 18, 2025
1,588 words in the original blog post.
The Arcee Foundation Models introduce a new family of generative AI models designed for enterprise applications, with the initial release being the AFM-4.5B, a 4.5-billion-parameter model noted for its accuracy, compliance, and cost-efficiency. Unlike larger language models that are costly and pose data privacy and customization challenges, AFM-4.5B is built to deliver high performance at lower costs and can be deployed across various environments, including smartphones, edge devices, and cloud platforms. The model is developed with an open architecture, allowing customizable deployment without compromising on compliance or sovereignty, and is trained on a clean dataset of almost 7 trillion tokens to minimize intellectual property risks. It supports multiple languages, and its modular training allows for easy language expansion and domain-specific customizations. Licensing options are flexible, with non-commercial use available under a CC-BY-NC license, and commercial deployment supported through direct or white-label licensing. The AFM model family aims to provide scalable, industry-specific solutions with rapid deployment capabilities, offering enterprises a foundation model that balances performance, cost efficiency, and compliance.
Jun 18, 2025
973 words in the original blog post.
Arcee AI has developed an innovative solution to address the challenge of enabling different small language models (SLMs) to work together despite having distinct vocabularies, a problem traditionally solved through costly retraining. Their research, titled "Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit," introduces a method known as tokenizer transplantation, which allows models to convert between different vocabularies without retraining. This approach identifies common ground between model vocabularies and applies familiar patterns to a target model's vocabulary space, preserving performance and significantly reducing costs and time. The method has shown impressive results in maintaining model performance across tasks and enabling cross-language compatibility, with applications in knowledge distillation, speculative decoding, model merging, domain adaptation, and cross-language model development. This breakthrough, available through their open-source MergeKit library, exemplifies Arcee AI's commitment to enhancing AI development by making it faster, cost-effective, and more collaborative.
Jun 10, 2025
803 words in the original blog post.
The Edge IQ Retail Assistant exemplifies the transformative potential of AI-driven solutions in the retail sector by leveraging CPU-based edge computing to enhance store operations and customer interactions. It operates without GPUs, utilizing Intel Xeon 6 CPUs within a Cisco UCS server to process data locally, thereby addressing key retail challenges such as latency, resilience, data privacy, bandwidth efficiency, and cost predictability. The system integrates open-source small language models and real-time data analytics, allowing store associates to access information via voice or text on customer traffic and inventory from platforms like WaitTime and Chooch. This approach reduces network dependency, ensures continuous functionality during internet outages, and keeps sensitive data within store infrastructure. By using Intel's OpenVINO toolkit, the assistant optimizes AI models for CPU execution, demonstrating the significant advancements in AI optimization and the capabilities of modern server processors. As a result, the solution offers enhanced customer experiences, improved operational efficiency, and cost savings, positioning retailers to remain competitive in an evolving landscape.
Jun 07, 2025
1,261 words in the original blog post.
Madeline & Co., an AI-powered strategy and design platform, faced challenges with existing large language models due to high costs and lack of domain-specific reasoning, leading them to collaborate with Arcee AI to develop a custom reasoning model named Madeline-s1. The collaboration involved creating a 60 million token dataset and employing Continuous Pre-Training, model merging, and Supervised Fine-Tuning to enhance the model's reasoning capabilities and domain-specific accuracy. This process included rigorous behavioral analysis and synthetic data generation to address identified weaknesses. The resulting 32-billion-parameter model excelled in generating actionable insights and was evaluated through blind human preference tests and industry-standard benchmarks, outperforming general models in business, strategy, design, and storytelling domains. Madeline-s1's success highlights the potential of combining proprietary data with specialized AI training techniques to create tailored solutions, ultimately leading to its integration into Madeline & Co.'s core offerings.
Jun 06, 2025
838 words in the original blog post.