Home / Companies / Mixedbread / Blog / August 2024

August 2024 Summaries

2 posts from Mixedbread

Filter
Month: Year:
Post Summaries Back to Blog
Baguetter is an open-source, Python-based retrieval testing framework designed to unify sparse, dense, and hybrid search experimentation through a single extensible interface. Created to address limitations in existing tools encountered during work on BMX and hybrid-search fusion algorithms, it supports lexical retrieval, semantic embeddings, reranking, and embedding quantization while aiming to remain lightweight and performant. Forked from the retriv project, Baguetter adds keyword-search implementations including BM25S and BMX, and uses USearch and Faiss for dense retrieval. The framework can be installed with pip, indexed with documents, and used to compare retrieval approaches through a common API. For evaluation, it relies on ranx and supports metrics including nDCG, precision, mean reciprocal rank, and recall, while providing Hugging Face dataset wrappers for benchmarks such as MTEB and tools to evaluate and save results from alternative index implementations.
Aug 23, 2024 925 words in the original blog post.
Mixedbread and Hong Kong Polytechnic University researchers introduced BMX, an open-source lexical search algorithm available through the Baguetter library that aims to improve on BM25 while retaining the efficiency and generalization strengths of keyword-based retrieval. BMX adds entropy-weighted similarity to emphasize informative, less common query terms and uses weighted query augmentation to incorporate semantic variations in a single retrieval process without separate reranking. Evaluations on BEIR found BMX outperformed BM25 variants on 11 of 15 datasets, while BMX with weighted query augmentation achieved the strongest average result on the reasoning-intensive BRIGHT benchmark, surpassing the compared lexical, embedding, and proprietary models. Additional tests in Chinese, Japanese, Korean, German, and French showed consistent gains over BM25. The researchers argue that BMX can improve search quality and downstream NLP pipelines without the large training datasets or computational costs commonly associated with embedding-based semantic search.
Aug 12, 2024 1,593 words in the original blog post.