Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Rikiya Takehi, Benjamin ClaviƩ, Sean Lee
Word Count
1,762
Company Posts That Month
4
Language
English
Hacker News Points
2
Post removed?
No
Summary

MXBAI introduced the open-source mxbai-edge-colbert-v0 late-interaction retrieval models in 17-million- and 32-million-parameter versions, designed as compact, reproducible baselines for retrieval research and edge deployment. Built on small Ettin/ModernBERT-style encoders, the models were trained through weakly supervised contrastive pretraining, supervised retrieval fine-tuning with hard negatives, and Stella-inspired knowledge distillation before ColBERT-specific optimization and ablation studies. Benchmark results indicate that the 17M version outperforms ColBERTv2 despite using a 48-dimensional projection, while both variants perform strongly on BEIR and LongEmbed, including long-context retrieval. Their low parameter counts, small projection dimensions, Flash Attention 2 support, and built-in unpadding reduce memory and CPU requirements, making them suitable for embedding or reranking documents on modest hardware. Both checkpoints are available through Hugging Face and supported by PyLate, with the developers planning future updates to their edge-focused retrieval models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 6 1,855 367 153 +5%
AI Model Fine-tuning 1 546 132 69 +43%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.