Home / Companies / Arcee AI / Blog / Post Details
Content Deep Dive

Arcee-SuperNova: Training Pipeline and Model Composition

Blog post from Arcee AI

Post Details
Company
Date Published
Author
Lucas Atkins and Fernando Fernandes
Word Count
1,353
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

Arcee-Llama-3.1-SuperNova, a language model intended to replace larger proprietary models, focuses on instruction-following and human preference alignment. The development involved distilling Llama-3.1-405B-Instruct into a more manageable 70B version, overcoming computational challenges through logits compression, which reduced the dataset size significantly. Training utilized Fully Sharded Data Parallel (FSDP) and Spectrum with EvolKit for parallel processing, enhancing the model's performance on benchmarks like reasoning and math. Despite outperforming models like GPT-4 in some areas, SuperNova underperforms in specific benchmarks such as GPQA and MUSR. The model's training incorporated several novel techniques, such as checkpoint merging strategies and Direct Preference Optimization (DPO), resulting in a model that offers precise control and performance reliability for business applications. A smaller 8B variant, SuperNova-Lite, was also developed, maintaining key capabilities while providing a lightweight alternative. Future efforts aim to enhance its robustness and expand its applicability while ensuring full control and security for users.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.