Open Source Gets DE-licious: Mixedbread x deepset German/English Embeddings
Blog post from Mixedbread
Deepset and Mixedbread have released deepset-mxbai-embed-de-large-v1, an open-source German/English embedding model designed primarily for retrieval tasks and based on multilingual-e5-large. Fine-tuned on more than 30 million German data pairs using AnglE loss and a combination of full fine-tuning and LoRA, the model aims to address the limited quality of German-focused embedding tools, which are often overshadowed by English-oriented models. In reported benchmarks, it achieved an NDCG@10 score of 51.7, surpassing other open-source German models and approaching Cohere Multilingual v3, while a legal-data case study reported a MAP@10 of 90.25, exceeding a domain-specific legal embedding model. The model supports binary quantization and Matryoshka representation learning, features intended to reduce vector storage and computing requirements while retaining much of its retrieval quality, with binary quantization reportedly preserving 91.8% performance at 32 times greater efficiency and reduced dimensions offering further size-performance trade-offs. The release is available through Mixedbread and can be used with APIs and common embedding frameworks, while the collaborators invite community feedback and acknowledge NVIDIA’s donated DGX A100 computing resources for training and evaluation.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 32 | 1,704 | 240 | 102 | -4% |
| LLM | 4 | 4,537 | 421 | 147 | +51% |
| AI Model Fine-tuning | 3 | 1,029 | 157 | 78 | +15% |
| RAG | 3 | 1,801 | 200 | 85 | +50% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.