Arcee-SuperNova: Training Pipeline and Model Composition
Blog post from Arcee AI
Arcee-Llama-3.1-SuperNova, a language model intended to replace larger proprietary models, focuses on instruction-following and human preference alignment. The development involved distilling Llama-3.1-405B-Instruct into a more manageable 70B version, overcoming computational challenges through logits compression, which reduced the dataset size significantly. Training utilized Fully Sharded Data Parallel (FSDP) and Spectrum with EvolKit for parallel processing, enhancing the model's performance on benchmarks like reasoning and math. Despite outperforming models like GPT-4 in some areas, SuperNova underperforms in specific benchmarks such as GPQA and MUSR. The model's training incorporated several novel techniques, such as checkpoint merging strategies and Direct Preference Optimization (DPO), resulting in a model that offers precise control and performance reliability for business applications. A smaller 8B variant, SuperNova-Lite, was also developed, maintaining key capabilities while providing a lightweight alternative. Future efforts aim to enhance its robustness and expand its applicability while ensuring full control and security for users.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.