Candidate Depth: How Much Retrieval Is Enough?
Blog post from Qdrant
Candidate depth determines how many retrieved items reach later ranking stages, making it valuable only when fusion, reranking, or payload-based rescoring can benefit from additional candidates. Qdrant’s analysis recommends establishing a labeled baseline, testing limits around 100 to 200, and comparing the candidate set’s best possible nDCG@10 with current ranking performance to distinguish retrieval limitations from ranking shortcomings. Across several datasets, deeper retrieval substantially increased the potential score available to downstream rankers while default reciprocal rank fusion often changed little, and higher limits increased latency by 37% to 43% in single-shard tests. For dense search, hnsw_ef should be increased only while recall against exact search continues improving, since wider HNSW traversal can otherwise add latency without meaningful relevance gains; graph construction settings, filtering behavior, and collection scale can also constrain recall. When memory is the main concern, reducing candidate depth lowers query work but does not reduce vector storage, whereas int8 scalar quantization can cut vector size to one quarter with minimal observed impact on hybrid ranking, particularly when rescoring is enabled. The appropriate next tuning step depends on whether relevant documents are already present but poorly ranked, in which case fusion or reranking should be tested, or absent from the candidate set, in which case retrieval settings should be improved.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.