March 2024 Summaries
3 posts from Mixedbread
Filter
Month:
Year:
Post Summaries
Back to Blog
Mixedbread AI introduces mxbai-colbert-large-v1, an Apache 2.0-licensed ColBERT model available on Hugging Face that is designed for retrieval-augmented generation and reranking. ColBERT bridges standard embedding search and compute-intensive cross-encoders by encoding query and document tokens separately, then applying MaxSim late-interaction scoring to capture fine-grained relevance signals efficiently. Initialized from mxbai-embed-large-v1, which was trained on more than 700 million diverse samples, the model was further adapted using about 96 million samples assembled from cleaned web data. The company reports that, as of March 2024, the model achieved the highest average NDCG@10 score among compared ColBERT systems across 13 public BEIR reranking benchmarks and performed strongly on three tested retrieval tasks, while noting that full retrieval evaluation remained incomplete. It recommends using the model through the RAGatouille framework and continues to suggest its standard embedding model for primary retrieval use cases.
Mar 19, 2024
1,561 words in the original blog post.
Mixedbread AI announced mxbai-embed-large-v1, an Apache 2.0-licensed English embedding model available on Hugging Face and designed for retrieval-augmented generation, semantic search, and related applications. Embeddings convert documents into vector representations that can be searched to retrieve relevant internal or external information for generative models. The company says the model can be integrated into existing retrieval pipelines through local hosting or an upcoming API, with a recommended query prompt for information-retrieval tasks. It was trained on more than 700 million contrastive pairs and tuned on over 30 million triplets from internally constructed web data, while excluding potential overlap with most MTEB benchmark test data. On the 56-dataset Massive Text Embedding Benchmark, the model reported a 64.68 average score, placing it ahead of similarly sized open-source models and slightly above OpenAI’s text-embedding-3-large overall, though performance varied by task. Mixedbread AI chose not to emphasize long-context embeddings, arguing that single vectors cannot reliably capture multiple unrelated topics in lengthy documents, and said it is developing Matryoshka-compatible and preference-improved versions while inviting community feedback.
Mar 08, 2024
1,433 words in the original blog post.
Mixedbread AI has released mxbai-embed-2d-large-v1, an Apache 2.0-licensed embedding model available on Hugging Face that introduces 2D Matryoshka representation learning, allowing users to reduce both the model’s layer count and embedding dimensionality. Designed for retrieval-augmented generation and semantic search applications, the approach aims to provide configurable tradeoffs among inference speed, memory use, storage requirements, retrieval efficiency, and accuracy by deriving smaller usable models from a single 24-layer model. The model was contrastively pretrained on more than 700 million text pairs and fine-tuned on over 30 million triplets, with training intended to avoid overlap with most MTEB benchmark data. Its reported MTEB score of 63.25 is competitive with several established embedding models, though the authors note that it may trail some larger state-of-the-art systems. Benchmark results indicate that dimensional truncation can preserve competitive performance in tasks such as semantic textual similarity and retrieval, while reducing the model to 13 layers, roughly half its original depth, retains about 75% performance on SciFact and more than 85% on semantic textual similarity; the company invites feedback as it continues developing the technique.
Mar 04, 2024
1,858 words in the original blog post.