Reasoning Tokens And Memory: 8x Cost Reduction Test
Blog post from Mem0
An experiment evaluated whether a Mem0 memory layer could reduce the cost of high-effort reasoning models without materially reducing answer quality, comparing full conversation-history prompts with retrieved, distilled memories for GPT-5 mini and Gemini 3.6 Flash. Using five LoCoMo-inspired conversations containing relevant facts, plausible distractors, and recall questions, the test employed Claude Haiku 4.5 as an independent, deterministic answer-quality judge and kept reasoning effort consistent across conditions. Memory reduced input tokens modestly but produced much larger reductions in output or reasoning tokens: GPT-5 mini fell from an average of 9,651 to 1,136 output tokens and from $0.0193 to $0.0023 per question, while Gemini fell from 3,228 to 660 output tokens and from $0.0245 to $0.0052. The full-history condition answered all ten cases correctly, whereas the memory condition answered nine, with the single error involving GPT-5 mini selecting an outdated value despite memory identifying a later value as an override. The findings suggest that memory can reduce repeated context processing and reasoning costs by roughly five to eight times, although effectiveness depends on retrieval quality and on how a deployed model interprets retrieved memories.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.