64 bytes per embedding, yee-haw đź¤
Blog post from Mixedbread
Mixedbread introduces Binary MRL, a compression approach that combines Matryoshka Representation Learning, which prioritizes important information in earlier embedding dimensions, with binary quantization, which converts float32 values into one-bit representations. Designed for its mxbai-embed-large-v1 model, the method aims to reduce the memory, storage, latency, and cloud costs associated with large-scale vector search while retaining most retrieval quality. On MTEB’s 13 BEIR retrieval datasets, 512-dimensional binary embeddings retained about 90.8% of the performance of full 1024-dimensional float32 vectors while reducing vector size from 4,096 bytes to 64 bytes, a 64-fold efficiency gain. The reported savings could substantially reduce infrastructure costs for indexes containing hundreds of millions or billions of embeddings, potentially making additional semantic-search and retrieval-augmented generation applications economically practical. Binary MRL is available through Mixedbread’s API and Sentence Transformers, though its full benefits require vector databases that support binary embeddings; Vespa is identified as an early supporter.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 41 | 2,722 | 279 | 102 | +43% |
| RAG | 3 | 1,867 | 232 | 78 | +54% |
| LLM | 2 | 3,669 | 412 | 154 | +40% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.