Home / Companies / Octen / Blog / Post Details
Content Deep Dive

Octen Series: Optimizing Embedding Models to #1 on RTEB Leaderboard

Blog post from Octen

Post Details
Company
Date Published
Author
Octen Algorithm Team
Word Count
2,950
Company Posts That Month
1
Language
中文
Hacker News Points
-
Post removed?
No
Summary

RTEB (Retrieval Embedding Benchmark) is a new benchmark developed by MTEB to evaluate retrieval capabilities in real-world industry scenarios, addressing issues like model overfitting on public datasets. Focusing on domains such as legal, finance, healthcare, and code, RTEB utilizes a hybrid evaluation strategy with both open and private datasets to measure true generalization capabilities. The Octen series models, particularly the Octen-Embedding-8B, achieved first place on the RTEB leaderboard by leveraging systematic technical optimizations, such as domain-specific synthetic data, data optimization strategies like Hard Negative sampling, and training efficiency improvements through methods like LoRA fine-tuning. Despite challenges regarding evaluation fairness due to unequal access to private datasets, Octen models demonstrated their generalization capabilities and maintained high performance across both public and private datasets. The series has been open-sourced on the Hugging Face platform, promoting technological progress and community contribution in the field of retrieval and embedding technologies.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 61 2,057 332 133 +28%
AI Model Fine-tuning 8 593 154 74 -13%
Real-time 1 6,429 1,407 265 -24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.