Introducing DolphinBench: Mapping The Pareto Frontier Of Agent Memory
Blog post from Mem0
DolphinBench is an open benchmark for evaluating long-term memory in AI agents through simulated real-world actions rather than conversational fact-recall questions, emphasizing whether agents can recognize when memory is needed and correctly apply stored information. It includes 600 tool-using tasks across three personas with multi-year histories totaling roughly 500,000 tokens each, requiring agents to interact with simulated workplace applications such as Gmail, Notion, GitHub, Discord, calendars, and CRMs. The benchmark measures accuracy together with total cost and median task latency, reflecting the trade-offs between memory performance and practical deployment constraints. Each task is certified through runs in which an agent must succeed when supplied with relevant history and fail when that history is removed, helping exclude unsolvable or improperly graded tests. Initial results show that the highest-performing tested configuration, Hermes with GPT-5.6-Luna and Mem0, achieved 70.67% memory accuracy, while results varied substantially by model, harness, cost, and memory system. DolphinBench is publicly available with a dataset, repository, evaluation harness, and leaderboard, and its developers plan to expand it with longer histories, additional personas, more complex tasks, and ingestion-testing patterns closer to production use.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.