Why Agents Need Custom Evals and How to Create Them [Testμ 2026]
Blog post from TestMu AI
Haritha Sreedharan Nair of Oqoqo argued at Testμ Conf 2026 that custom agent benchmarks should function like product-specific regression suites, testing whether agents can complete real user workflows across actual documentation, tools, permissions, APIs, integrations, and system states rather than merely producing correct-looking outputs. Her central example described an evaluation of a vector database that passed because the agent quietly switched to SQLite, illustrating why trajectories—including tool calls, retries, latency, bypasses, and recovery behavior—should be graded alongside final results. She distinguished custom benchmarks from public model leaderboards, which remain useful for comparing models but generally lack the product context needed to reveal real-world failures. Her recommendations included designing tasks from users’ jobs rather than API surfaces or tutorial-style prompts, using fair rubrics that assess both outputs and paths taken, running repeated trials for probabilistic systems, classifying failures into actionable categories, and using isolated test accounts with representative data instead of mocks. She also noted practical obstacles such as rate limits, credential boundaries, stale benchmarks, and uncertainty over whether companies view agents as users or merely integrations, before demonstrating Oqoqo’s benchmark platform and emphasizing that teams should begin by defining their most important user workflows and success criteria.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MCP | 7 | 2,241 | 148 | 72 | -74% |
| Vector Search | 3 | 265 | 57 | 33 | -89% |
| Loop engineering | 1 | 16 | 8 | 7 | -77% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.