Home / Companies / Arcee AI / Blog / Post Details
Content Deep Dive

Extending AFM-4.5B to 64k Context Length

Blog post from Arcee AI

Post Details
Company
Date Published
Author
Charles Goddard
Word Count
2,752
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

Arcee has unveiled its new foundation model, AFM-4.5B, which boasts an extended context length from 4k to 64k through a rigorous and experimental training process. This model, intended for both short and long-context tasks, is a product of various innovative techniques including model merging, distillation, and context extension strategies derived from existing research. Initial evaluations reveal promising results, though these are preliminary and subject to change as further training continues. AFM-4.5B's development involved a series of experiments, leveraging methods like YaRN positional embeddings, ProLong training, and model merging to enhance performance while maintaining short-context task efficiency. The model's final form achieved notable improvements in benchmarks such as MMLU and Big Bench Hard, demonstrating a balance between maintaining short-context accuracy and extending long-context capabilities. The process highlights the effectiveness of linear averaging and distillation in refining model performance, suggesting that these methods can scale to larger models as well, as evidenced by successful trials on the GLM-32B base model.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 7 1,525 253 110 -6%
LLM 1 3,482 526 172 -8%
Reinforcement learning 1 114 37 24 -27%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.