Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Long context retrieval models with Monarch Mixer

Blog post from Together AI

Post Details
Company
Date Published
Author
Jon Saad-Falcon, Dan Fu, Simran Arora
Word Count
2,583
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

The text discusses the development of long-context retrieval models using Monarch Mixer, a recent model family that aims to improve the scaling properties of Transformers along two axes – sequence length and model dimension. The authors release a preview of several models, including long-context versions of M2-BERT up to 32K context length and embedding versions fine-tuned for long-context retrieval. They also introduce a new benchmark called LoCo, which is designed to evaluate the performance of long-context retrieval models on tasks with long documents. The authors report promising results, demonstrating that their long-context M2-BERT models can outperform much larger models on this benchmark, suggesting that long-context models are beneficial for retrieval. The authors also discuss challenges in training long-context models, including adapting the BERT pretraining pipeline and fine-tuning the model using a suitable loss function. They propose a new loss function called orthogonal loss, which pushes the cosine similarity of positive pairs to 1 and the cosine similarity of negative pairs to 0.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 18 1,728 228 84 +63%
AI Model Fine-tuning 3 444 125 69 +22%
RAG 2 1,418 170 60 +93%
AI Guardrails 1 88 50 26 +38%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.