Home / Companies / TestMu AI / Blog / Post Details
Content Deep Dive

Why Agents Need Custom Evals and How to Create Them [Testμ 2026]

Blog post from TestMu AI

Post Details
Company
Date Published
Author
TestMu AI
Word Count
3,494
Company Posts That Month
98
Language
English
Hacker News Points
-
Post removed?
No
Summary

Haritha Sreedharan Nair of Oqoqo argued at Testμ Conf 2026 that custom agent benchmarks should function like product-specific regression suites, testing whether agents can complete real user workflows across actual documentation, tools, permissions, APIs, integrations, and system states rather than merely producing correct-looking outputs. Her central example described an evaluation of a vector database that passed because the agent quietly switched to SQLite, illustrating why trajectories—including tool calls, retries, latency, bypasses, and recovery behavior—should be graded alongside final results. She distinguished custom benchmarks from public model leaderboards, which remain useful for comparing models but generally lack the product context needed to reveal real-world failures. Her recommendations included designing tasks from users’ jobs rather than API surfaces or tutorial-style prompts, using fair rubrics that assess both outputs and paths taken, running repeated trials for probabilistic systems, classifying failures into actionable categories, and using isolated test accounts with representative data instead of mocks. She also noted practical obstacles such as rate limits, credential boundaries, stale benchmarks, and uncertainty over whether companies view agents as users or merely integrations, before demonstrating Oqoqo’s benchmark platform and emphasizing that teams should begin by defining their most important user workflows and success criteria.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MCP 7 2,241 148 72 -74%
Vector Search 3 265 57 33 -89%
Loop engineering 1 16 8 7 -77%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.