RCP-nDCG@10: A more complete way to measure retrieval relevance
Blog post from Cohere
Cohere introduces Rubric-Calibrated Preferences nDCG@10 (RCP-nDCG@10), a retrieval-evaluation method intended to address limitations in conventional nDCG, which depends on incomplete pre-existing relevance labels and may fail to reward newly retrieved but useful documents. While Recall@k, MRR, and nDCG measure different dimensions of search quality, Cohere argues that sparse benchmark labels increasingly constrain evaluation as retrieval models improve; in its human study, 28% of documents benchmarked as irrelevant were considered useful by reviewers. RCP-nDCG@10 uses a calibrated AI judge that applies consistent yes-or-no relevance rubrics to every retrieved document and compares documents in groups, combining rubric scores and pairwise preferences to create cross-query relevance scores. In a blind evaluation involving 46 annotators and 289 system comparisons, the methodology selected the human-preferred system 77% of the time, compared with 52% for traditional nDCG in a deliberately disagreement-heavy sample. Cohere says it has optimized its forthcoming fifth-generation Embed and Rerank models using RCP-nDCG@10, prioritizing alignment with perceived search usefulness rather than legacy benchmark performance alone.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 747 | 162 | 79 | -85% |
| Vector Search | 2 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.