Evaluating Agents Against What They Are Actually Supposed to Do [Testμ 2026]
Blog post from TestMu AI
Microsoft’s Francesca Lazzeri argued at Testμ Conf 2026 that generic AI benchmarks can overlook critical agent failures, such as violating refund policies, approval limits, or instructions embedded in tool results, even when scores for helpfulness, groundedness, and safety are high. She presented ASSERT, Microsoft’s open-source Adaptive Specification-driven Scoring for Evaluation and Regression Testing framework, which turns plain-language requirements about what an agent must and must not do into behavior taxonomies, generated tests, and scored evaluation results. Her four-layer evaluation loop begins with written specifications, adds shared baseline metrics such as groundedness, retrieval quality, relevance, fluency, and safety, incorporates domain-specific measures including task adherence and tool-call accuracy, and uses production observability to identify unexpected failures that become future tests. Using a support assistant with refund capabilities as an example, she emphasized that evaluating only final answers can obscure unsafe or noncompliant action sequences. The observability layer also links quality with token costs, workflow outcomes, revenue, and controlled or causal experimentation to assess business impact. Lazzeri concluded that reliable agent evaluation requires both standardized measures for comparison across systems and custom metrics developed with product teams, domain experts, and users.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 7 | 472 | 102 | 54 | -85% |
| AI Agents | 5 | 931 | 231 | 103 | -84% |
| Real-time | 2 | 649 | 155 | 80 | -85% |
| AI Guardrails | 1 | 35 | 22 | 12 | -94% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.