Redefining Quality Leadership in an Agentic World [Testμ 2026]
Blog post from TestMu AI
A Testμ Conf 2026 session led by Salesforce engineering manager Sobhitha Neelanath argued that conventional deterministic testing metrics, such as a 94.2% pass rate, can obscure serious failures in probabilistic agentic systems, illustrated by a recursive refund loop that issued $400,000 before alerts activated. She proposed measuring trust through an equation combining alignment and predictability while minimizing blast radius, supported by indicators including conversational accuracy, intent continuity, output variance, and mitigation capability. The session categorized risks by agent autonomy, from context hallucinations in assistants to cascading workflow failures in autonomous agents, and recommended progressively stronger controls such as grounded assertions, human review gates, and transaction isolation. Neelanath also urged leaders to replace broad test coverage and productivity counts with impact-based regression, shorter feedback cycles, defect leakage measures, and commercial reliability outcomes, citing a reported 78% gap between executive expectations and actual agentic performance. Her team’s use of “chaos audits,” which rewards engineers for deliberately inducing failures and improving safeguards, was associated with reported reductions in production loops and burnout. She described an evolving quality-engineering path from script automator to evaluator, trust architect, and governance-focused guardian, requiring semantic evaluation skills, statistical rigor, and compliance awareness while retaining human involvement for complex customer workflows and system design decisions.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Multi-agent systems | 3 | 41 | 24 | 19 | -91% |
| AI Agents | 2 | 931 | 231 | 103 | -84% |
| Voice AI | 1 | 324 | 41 | 16 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.