Dense Retrievers Know More Than They Can Express
Blog post from Mixedbread
Neural retrieval models are argued to be limited less by the information they learn than by the scoring operators used to rank documents, with single-vector cosine similarity restricting what dense representations can express compared with late-interaction MaxSim approaches. The authors use sparse autoencoders (SAEs) to extract sparse “Latent Terms” from dense retriever activations without additional retrieval training, finding that these features form a roughly Zipfian distribution similar to natural-language vocabularies and include lexical, narrow semantic, and broad topical concepts. Because these sparse features resemble lexical terms, they can be indexed and ranked using BM25, producing retrieval performance that is competitive with or better than the originating single-vector models and, in some evaluations, comparable to SPLADE. On the LIMIT benchmark, Latent Terms substantially improve a dense model’s ability to retrieve documents requiring fine-grained attribute matching, supporting the view that dense retrievers encode relevance information not accessible through their usual scoring method. Experiments also indicate that this retrieval-ready sparse structure arises from retrieval-focused training rather than generic pretrained language-model representations, raising questions about better ways to extract and train such latent retrieval vocabularies.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 8 | 1,918 | 398 | 137 | -21% |
| LLM | 4 | 6,292 | 1,205 | 252 | -36% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.