A Delicious Free Lunch: Better Projections Improve ColBERT
Blog post from Mixedbread
ColBERT-style late-interaction retrieval models represent queries and documents as token-level vectors and rank documents using MaxSim, which retains only each query token’s highest similarity to a document token. The authors argue that this scoring mechanism creates sparse gradient flow during training, potentially updating only a small fraction of document-token representations, and that its preference for highly discriminative similarity peaks makes the conventional single-layer projection head suboptimal. They propose deeper projection heads with options including intermediate upscaling, residual connections, nonlinear activations, and GLU gating, aiming to preserve and sharpen diverse token representations more effectively. Experiments using modified PyLate training tools across hundreds of controlled model runs found that most projection variants significantly improved retrieval performance, with the strongest configurations averaging roughly two NDCG@10 points, or nearly 4% relative improvement, while adding only a small number of parameters and little practical cost.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| RAG | 2 | 1,142 | 236 | 104 | -1% |
| Vector Search | 1 | 1,855 | 367 | 153 | +5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.