Home / Companies / Mixedbread / Blog / October 2024

October 2024 Summaries

1 posts from Mixedbread

Filter
Month: Year:
Post Summaries Back to Blog
Mixedbread introduces mxbai-embed-xsmall-v1, an Apache 2.0-licensed English embedding model available on Hugging Face that targets retrieval applications under limited computational resources. Built from sentence-transformers/all-MiniLM-L6-v2, it contains 22.7 million parameters, produces 384-dimensional embeddings, supports contexts up to 4,096 tokens, and is designed for search, recommendations, clustering, classification, and retrieval-augmented generation. Benchmark results show modest improvements over its base model on average MTEB retrieval scores and larger gains on several long-context evaluations. The model supports Matryoshka representation learning, which retains useful performance at reduced embedding dimensions, and binary quantization, which can reduce storage by up to 32 times and computation by up to 40 times with a limited benchmark performance decrease. It was trained with AnglE loss and Espresso techniques to prioritize English semantic retrieval, aiming to provide faster inference, lower memory use, and lower deployment costs for large-scale or resource-constrained systems.
Oct 14, 2024 1,236 words in the original blog post.