Home / Companies / Mixedbread / Blog / Post Details
Content Deep Dive

Open Source Strikes Bread - New Fluffy Embedding Model

Blog post from Mixedbread

Post Details
Company
Date Published
Author
Sean Lee, Aamir Shakir, Darius Koenig, Julius Lipp
Word Count
1,433
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Mixedbread AI announced mxbai-embed-large-v1, an Apache 2.0-licensed English embedding model available on Hugging Face and designed for retrieval-augmented generation, semantic search, and related applications. Embeddings convert documents into vector representations that can be searched to retrieve relevant internal or external information for generative models. The company says the model can be integrated into existing retrieval pipelines through local hosting or an upcoming API, with a recommended query prompt for information-retrieval tasks. It was trained on more than 700 million contrastive pairs and tuned on over 30 million triplets from internally constructed web data, while excluding potential overlap with most MTEB benchmark test data. On the 56-dataset Massive Text Embedding Benchmark, the model reported a 64.68 average score, placing it ahead of similarly sized open-source models and slightly above OpenAI’s text-embedding-3-large overall, though performance varied by task. Mixedbread AI chose not to emphasize long-context embeddings, arguing that single vectors cannot reliably capture multiple unrelated topics in lengthy documents, and said it is developing Matryoshka-compatible and preference-improved versions while inviting community feedback.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 35 1,909 252 81 -13%
RAG 5 1,215 181 58 +4%
AI Guardrails 1 112 45 22 +2%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.