Home / Companies / Supermemory / Blog / Post Details
Content Deep Dive

2 Approaches For Extending Context Windows in LLMs

Blog post from Supermemory

Post Details
Company
Date Published
Author
Naman Bansal
Word Count
2,094
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

Transformer-based language models are limited by finite context windows and the quadratic computational cost of self-attention, which can force applications to truncate or repeatedly summarize long inputs. Two approaches address this issue: semantic compression for very long static documents and Infinite Chat for extended conversations. Semantic compression divides normalized text into sentence-sized blocks, creates MiniLM embeddings and a similarity graph, uses spectral clustering to identify coherent topics, summarizes clusters in parallel with BART, and reassembles the results in original order, reportedly achieving roughly 6:1 compression while retaining strong retrieval performance. The guide provides Python setup and implementation examples for document loading, token-aware chunking, similarity calculation, clustering, parallel summarization, and prompting an LLM with the compressed output. Infinite Chat, provided through Supermemory’s proxy for OpenAI-compatible APIs, stores conversation chunks in an embedding index and dynamically reconstructs prompts by ranking historical material according to relevance and recency within a fixed token budget. Together, these methods aim to reduce token use and preserve useful context for oversized documents and long-running chats without changing an underlying model’s architecture.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 8 4,922 763 224 +11%
Vector Search 4 2,058 362 133 +24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.