Home / Companies / Mixedbread / Blog / July 2024

July 2024 Summaries

1 posts from Mixedbread

Filter
Month: Year:
Post Summaries Back to Blog
Deepset and Mixedbread have released deepset-mxbai-embed-de-large-v1, an open-source German/English embedding model designed primarily for retrieval tasks and based on multilingual-e5-large. Fine-tuned on more than 30 million German data pairs using AnglE loss and a combination of full fine-tuning and LoRA, the model aims to address the limited quality of German-focused embedding tools, which are often overshadowed by English-oriented models. In reported benchmarks, it achieved an NDCG@10 score of 51.7, surpassing other open-source German models and approaching Cohere Multilingual v3, while a legal-data case study reported a MAP@10 of 90.25, exceeding a domain-specific legal embedding model. The model supports binary quantization and Matryoshka representation learning, features intended to reduce vector storage and computing requirements while retaining much of its retrieval quality, with binary quantization reportedly preserving 91.8% performance at 32 times greater efficiency and reduced dimensions offering further size-performance trade-offs. The release is available through Mixedbread and can be used with APIs and common embedding frameworks, while the collaborators invite community feedback and acknowledge NVIDIA’s donated DGX A100 computing resources for training and evaluation.
Jul 18, 2024 1,458 words in the original blog post.