Home / Companies / Weaviate / Blog / Post Details
Content Deep Dive

Late Chunking: Balancing Precision and Cost in Long Context Retrieval

Blog post from Weaviate

Post Details
Company
Date Published
Author
Charles Pierse, Connor Shorten, Akanksha Sharma
Word Count
2,517
Company Posts That Month
5
Language
English
Hacker News Points
2
Post removed?
No
Summary

JinaAI has introduced a new methodology called late chunking to aid in long-context retrieval for large documents. This approach aims to preserve contextual information across large documents by inverting the traditional order of embedding and chunking. Unlike naive chunking, which breaks up a document into chunks independently, or ColBERT, which requires significant storage capacity, late chunking maintains the contextual relationships between tokens across the entire document during the embedding process and only afterwards divides these contextually-rich embeddings into chunks. This method can help mitigate issues associated with very long documents, such as expensive LLM calls, increased latency, and a higher chance of hallucination. Late chunking offers a cost-effective path forward for users doing long context retrieval while preserving the contextual information that late interaction offers.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 27 3,701 290 90 +59%
RAG 8 1,966 260 82 -21%
LLM 2 4,030 486 147 +1%
Real-time 2 4,377 976 225 +49%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.