Home / Companies / Mem0 / Blog / Post Details
Content Deep Dive

Introducing DolphinBench: Mapping The Pareto Frontier Of Agent Memory

Blog post from Mem0

Post Details
Company
Date Published
Author
Deshraj Yadav
Word Count
1,121
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

DolphinBench is an open benchmark for evaluating long-term memory in AI agents through simulated real-world actions rather than conversational fact-recall questions, emphasizing whether agents can recognize when memory is needed and correctly apply stored information. It includes 600 tool-using tasks across three personas with multi-year histories totaling roughly 500,000 tokens each, requiring agents to interact with simulated workplace applications such as Gmail, Notion, GitHub, Discord, calendars, and CRMs. The benchmark measures accuracy together with total cost and median task latency, reflecting the trade-offs between memory performance and practical deployment constraints. Each task is certified through runs in which an agent must succeed when supplied with relevant history and fail when that history is removed, helping exclude unsolvable or improperly graded tests. Initial results show that the highest-performing tested configuration, Hermes with GPT-5.6-Luna and Mem0, achieved 70.67% memory accuracy, while results varied substantially by model, harness, cost, and memory system. DolphinBench is publicly available with a dataset, repository, evaluation harness, and leaderboard, and its developers plan to expand it with longer histories, additional personas, more complex tasks, and ingestion-testing patterns closer to production use.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.