2 Approaches For Extending Context Windows in LLMs
Blog post from Supermemory
Transformer-based language models are limited by finite context windows and the quadratic computational cost of self-attention, which can force applications to truncate or repeatedly summarize long inputs. Two approaches address this issue: semantic compression for very long static documents and Infinite Chat for extended conversations. Semantic compression divides normalized text into sentence-sized blocks, creates MiniLM embeddings and a similarity graph, uses spectral clustering to identify coherent topics, summarizes clusters in parallel with BART, and reassembles the results in original order, reportedly achieving roughly 6:1 compression while retaining strong retrieval performance. The guide provides Python setup and implementation examples for document loading, token-aware chunking, similarity calculation, clustering, parallel summarization, and prompting an LLM with the compressed output. Infinite Chat, provided through Supermemory’s proxy for OpenAI-compatible APIs, stores conversation chunks in an embedding index and dynamically reconstructs prompts by ranking historical material according to relevance and recency within a fixed token budget. Together, these methods aim to reduce token use and preserve useful context for oversized documents and long-running chats without changing an underlying model’s architecture.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 8 | 4,922 | 763 | 224 | +11% |
| Vector Search | 4 | 2,058 | 362 | 133 | +24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.