February 2024 Summaries
1 posts from Mixedbread
Filter
Month:
Year:
Post Summaries
Back to Blog
Mixedbread has released three Apache 2.0-licensed open-source reranking models—mxbai-rerank-xsmall-v1, base-v1, and large-v1—designed to improve search relevance while preserving existing keyword-search infrastructure such as Elasticsearch, OpenSearch, or Solr. Used as a second-stage step after initial retrieval, the models score and reorder candidate documents according to semantic relevance, offering organizations an alternative to fully migrating to embedding-based search. Trained on real-world queries and search-engine results labeled for relevance by a large language model, the family can be self-hosted or accessed through an upcoming API, with the base model positioned as a size-performance balance and the large model as the highest-accuracy option. On a subset of 11 BEIR datasets, Mixedbread reports that its models outperform lexical search and compare favorably with other rerankers, with the large model achieving 74.9 Accuracy@3 versus 66.4 for lexical search. The release includes local usage examples, integrations with common machine-learning tools, and an invitation for community feedback.
Feb 29, 2024
1,733 words in the original blog post.