July 2025 Summaries
8 posts from Supermemory
Filter
Month:
Year:
Post Summaries
Back to Blog
Supermemory has relaunched its consumer-focused product as a unified memory engine after initially operating as a manual “second brain” app with weak retention and later pivoting to B2B to solve underlying AI-memory challenges. The platform now automatically captures and updates information from AI tools such as ChatGPT, Claude, Cursor, and other MCP-compatible clients, while importing documents from Google Drive, Notion, and OneDrive. It organizes this material into searchable knowledge graphs that identify explicit and implicit relationships, supports project-based separation of personal and work information, and can deprioritize infrequently used memories to improve relevance and performance. Designed for nontechnical users, the app is currently available to a limited group through a waitlist, with manual uploads and MCP-based syncing options, including integrations for Claude Desktop and Cursor. Its free plan includes up to 10 memories and basic search, while an early paid tier offers expanded storage, connections, advanced search, and priority support for $9 per month; future plans include a browser extension for broader web-based memory capture.
Jul 25, 2025
1,034 words in the original blog post.
A tutorial describes building ContractGuard, an AI contract-compliance assistant that connects to Google Drive, synchronizes legal documents, and answers questions using document context and conversation memory. The application uses Supermemory for Google Drive integration, indexing, retrieval, and memory enhancement; OpenAI GPT-4o for responses; an Express and TypeScript backend; and a Vite, React, TypeScript, and Tailwind frontend. Backend routes initiate Google Drive OAuth, manually trigger synchronization, and process chat requests, with examples showing both automatic Supermemory context injection and a tool-based workflow in which the model decides whether to search connected documents before generating an answer. The frontend provides controls to connect and sync Drive, clear conversations, send messages, display Markdown-formatted assistant responses, and show loading states. After authorization and document upload, users can ask questions about their contracts, while the system retrieves relevant material and can identify obligations, risks, missing clauses, and recommended language without requiring manual file uploads or a custom retrieval pipeline.
Jul 20, 2025
3,548 words in the original blog post.
Supermemory, a portable memory engine for AI systems that integrates with multiple LLMs through APIs, SDKs, base URL swapping, and MCP, migrated from TigerData to PlanetScale after encountering limitations in branching, performance, observability, analytics, and escalating costs. Because its Infinite Chat product processes large volumes of vectors and user documents, the company required low-latency retrieval and reliable scaling, and PlanetScale supported a no-downtime migration using its proxy tooling and direct assistance. Following the move, Supermemory reports reducing monthly database costs from $900 to $90, increasing queries per second from 20 to 1,000, maintaining database p99 latency below 6 milliseconds, and gaining validated backups and improved query-performance visibility through PlanetScale Insights.
Jul 18, 2025
408 words in the original blog post.
Knowledge graph-based retrieval-augmented generation is presented as an alternative or complement to vector search for applications requiring explicit relationships, business-rule filtering, and explainable results. While embeddings support semantic similarity, knowledge graphs represent facts as entity-relation triples that can be queried through graph traversal and constraints, such as identifying suppliers of a product within a particular region. The tutorial demonstrates building a supply-chain question-answering system with Neo4j, Python, OpenAI-based triple extraction from procurement CSV records, Cypher queries, and template-based natural-language responses. It also outlines methods for evaluating graph quality, including coverage, accuracy, completeness, explainability, precision and recall, consistency checks, and manual audits. The discussion concludes that combining language models with graph retrieval can provide structured and reliable answers for domains containing interconnected data, while also mentioning Supermemory as a platform supporting vector and graph-style retrieval.
Jul 12, 2025
2,891 words in the original blog post.
Supermemory has expanded its Infinite Chat product into a context engineering system designed to preserve and manage AI conversation history more effectively across long-running, multi-step tasks. The release adds support for multiple LLM providers, including OpenAI, Gemini, and Anthropic, along with multimodal memory for images, audio, and structured data, awareness of tool calls and outputs, larger context handling, long-context retrieval-augmented generation, and automated summarization and pruning. Using a conversation ID, developers can avoid resending full message histories because the system reconstructs relevant context, tracks completed and pending work, and can draw on patterns from earlier interactions. The company also introduced the open-source llm-bridge package to normalize differences among provider APIs and enable integration with existing SDKs through a modified base URL. The updated architecture aims to reduce redundant context, token costs, and repeated tool actions while supplying models with the most relevant information for each subsequent decision.
Jul 09, 2025
783 words in the original blog post.
LLM-Bridge is an open-source TypeScript library created by Supermemory to address the complexity of supporting incompatible OpenAI, Anthropic, and Gemini API formats in services such as its Infinite Chat context-extension proxy. It defines a universal representation for LLM requests that supports lossless conversions between provider formats, preserving vendor-specific fields, multimodal content, tool definitions, and other capabilities while allowing middleware to inspect, modify, reroute, or enrich requests through a standardized normalize-process-emit workflow. The package also provides unified error translation, token-counting and other utilities, enabling use cases such as cross-provider routing, failover, load balancing, cost optimization, multimodal processing, and summarization with cheaper models without requiring users to change SDKs or client interfaces. Supermemory reports that it uses the library internally, that it has extensive Vitest coverage, and that the project was developed with assistance from Claude Code.
Jul 07, 2025
1,668 words in the original blog post.
Transformer-based language models are limited by finite context windows and the quadratic computational cost of self-attention, which can force applications to truncate or repeatedly summarize long inputs. Two approaches address this issue: semantic compression for very long static documents and Infinite Chat for extended conversations. Semantic compression divides normalized text into sentence-sized blocks, creates MiniLM embeddings and a similarity graph, uses spectral clustering to identify coherent topics, summarizes clusters in parallel with BART, and reassembles the results in original order, reportedly achieving roughly 6:1 compression while retaining strong retrieval performance. The guide provides Python setup and implementation examples for document loading, token-aware chunking, similarity calculation, clustering, parallel summarization, and prompting an LLM with the compressed output. Infinite Chat, provided through Supermemory’s proxy for OpenAI-compatible APIs, stores conversation chunks in an embedding index and dynamically reconstructs prompts by ranking historical material according to relevance and recency within a fixed token budget. Together, these methods aim to reduce token use and preserve useful context for oversized documents and long-running chats without changing an underlying model’s architecture.
Jul 04, 2025
2,094 words in the original blog post.
Presented through a fictional developer’s cost-overrun story, the post outlines practical approaches to reducing LLM expenses, including limiting context to task-relevant information, moving reusable instructions into system prompts, testing prompt variants for quality and token efficiency, caching static prompts and reusable personalization data, and requesting structured outputs to limit unnecessary generation. It emphasizes that targeted retrieval and context management can improve both performance and cost, while prompt A/B testing should evaluate correctness, latency, token use, and user satisfaction. The account also describes unsuccessful or risky approaches, such as self-hosting and fine-tuning open-source models, aggressive quantization that degrades quality, and serverless batch processing that can introduce latency and hidden costs. A central recommendation is to assess whether each task needs an LLM at all, since deterministic software or human review may be cheaper and more suitable. The author discloses that the developer narrative is fictional, while stating that the cited experts and opinions are real, and promotes Supermemory as a tool for optimizing long-context conversations.
Jul 01, 2025
2,267 words in the original blog post.