Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

Every Byte Matters: Introducing mxbai-embed-xsmall-v1

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Sean Lee, Julius Lipp, Rui Huang, Darius Koenig
Word Count
1,236
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Mixedbread introduces mxbai-embed-xsmall-v1, an Apache 2.0-licensed English embedding model available on Hugging Face that targets retrieval applications under limited computational resources. Built from sentence-transformers/all-MiniLM-L6-v2, it contains 22.7 million parameters, produces 384-dimensional embeddings, supports contexts up to 4,096 tokens, and is designed for search, recommendations, clustering, classification, and retrieval-augmented generation. Benchmark results show modest improvements over its base model on average MTEB retrieval scores and larger gains on several long-context evaluations. The model supports Matryoshka representation learning, which retains useful performance at reduced embedding dimensions, and binary quantization, which can reduce storage by up to 32 times and computation by up to 40 times with a limited benchmark performance decrease. It was trained with AnglE loss and Espresso techniques to prioritize English semantic retrieval, aiming to provide faster inference, lower memory use, and lower deployment costs for large-scale or resource-constrained systems.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 21 4,713 314 102 +27%
RAG 4 2,243 291 87 +14%
LLM 2 3,988 514 165 -1%
AI Guardrails 1 292 74 39 +93%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.