Octen Series: Optimizing Embedding Models to #1 on RTEB Leaderboard
Blog post from Octen
RTEB (Retrieval Embedding Benchmark) is a new benchmark developed by MTEB to evaluate retrieval capabilities in real-world industry scenarios, addressing issues like model overfitting on public datasets. Focusing on domains such as legal, finance, healthcare, and code, RTEB utilizes a hybrid evaluation strategy with both open and private datasets to measure true generalization capabilities. The Octen series models, particularly the Octen-Embedding-8B, achieved first place on the RTEB leaderboard by leveraging systematic technical optimizations, such as domain-specific synthetic data, data optimization strategies like Hard Negative sampling, and training efficiency improvements through methods like LoRA fine-tuning. Despite challenges regarding evaluation fairness due to unequal access to private datasets, Octen models demonstrated their generalization capabilities and maintained high performance across both public and private datasets. The series has been open-sourced on the Hugging Face platform, promoting technological progress and community contribution in the field of retrieval and embedding technologies.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 61 | 2,057 | 332 | 133 | +28% |
| AI Model Fine-tuning | 8 | 593 | 154 | 74 | -13% |
| Real-time | 1 | 6,429 | 1,407 | 265 | -24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.