Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

ColBERTus Maximus - Introducing mxbai-colbert-large-v1

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Sean Lee, Aamir Shakir, Darius Koenig, Julius Lipp
Word Count
1,561
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Mixedbread AI introduces mxbai-colbert-large-v1, an Apache 2.0-licensed ColBERT model available on Hugging Face that is designed for retrieval-augmented generation and reranking. ColBERT bridges standard embedding search and compute-intensive cross-encoders by encoding query and document tokens separately, then applying MaxSim late-interaction scoring to capture fine-grained relevance signals efficiently. Initialized from mxbai-embed-large-v1, which was trained on more than 700 million diverse samples, the model was further adapted using about 96 million samples assembled from cleaned web data. The company reports that, as of March 2024, the model achieved the highest average NDCG@10 score among compared ColBERT systems across 13 public BEIR reranking benchmarks and performed strongly on three tested retrieval tasks, while noting that full retrieval evaluation remained incomplete. It recommends using the model through the RAGatouille framework and continues to suggest its standard embedding model for primary retrieval use cases.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 10 1,909 252 81 -13%
RAG 4 1,215 181 58 +4%
AI Guardrails 1 112 45 22 +2%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.