Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

Fresh 2D-Matryoshka Embedding Model

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Sean Lee, Aamir Shakir, Julius Lipp, Darius Koenig
Word Count
1,858
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Mixedbread AI has released mxbai-embed-2d-large-v1, an Apache 2.0-licensed embedding model available on Hugging Face that introduces 2D Matryoshka representation learning, allowing users to reduce both the model’s layer count and embedding dimensionality. Designed for retrieval-augmented generation and semantic search applications, the approach aims to provide configurable tradeoffs among inference speed, memory use, storage requirements, retrieval efficiency, and accuracy by deriving smaller usable models from a single 24-layer model. The model was contrastively pretrained on more than 700 million text pairs and fine-tuned on over 30 million triplets, with training intended to avoid overlap with most MTEB benchmark data. Its reported MTEB score of 63.25 is competitive with several established embedding models, though the authors note that it may trail some larger state-of-the-art systems. Benchmark results indicate that dimensional truncation can preserve competitive performance in tasks such as semantic textual similarity and retrieval, while reducing the model to 13 layers, roughly half its original depth, retains about 75% performance on SciFact and more than 85% on semantic textual similarity; the company invites feedback as it continues developing the technique.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 47 1,909 252 81 -13%
RAG 2 1,215 181 58 +4%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.