Extending AFM-4.5B to 64k Context Length
Blog post from Arcee AI
Arcee has unveiled its new foundation model, AFM-4.5B, which boasts an extended context length from 4k to 64k through a rigorous and experimental training process. This model, intended for both short and long-context tasks, is a product of various innovative techniques including model merging, distillation, and context extension strategies derived from existing research. Initial evaluations reveal promising results, though these are preliminary and subject to change as further training continues. AFM-4.5B's development involved a series of experiments, leveraging methods like YaRN positional embeddings, ProLong training, and model merging to enhance performance while maintaining short-context task efficiency. The model's final form achieved notable improvements in benchmarks such as MMLU and Big Bench Hard, demonstrating a balance between maintaining short-context accuracy and extending long-context capabilities. The process highlights the effectiveness of linear averaging and distillation in refining model performance, suggesting that these methods can scale to larger models as well, as evidenced by successful trials on the GLM-32B base model.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 7 | 1,525 | 253 | 110 | -6% |
| LLM | 1 | 3,482 | 526 | 172 | -8% |
| Reinforcement learning | 1 | 114 | 37 | 24 | -27% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.