Salesforce's Four Phases of Agentic Testing [Testμ 2026]
Blog post from TestMu AI
Salesforce’s Service Cloud quality engineering team has adapted its testing strategy for AI agents through four phases: evaluating each stage of an agent’s reasoning and execution rather than only its final response, creating industry-specific environments based on real customer scenarios, using AI-driven agents to test complex voice interactions, and integrating customer-facing evaluation tools into Agentforce Studio. John Liang emphasized that deterministic assertions remain necessary alongside LLM judges, trace validation, metrics, and verification that an agent actually completed the task it claimed to perform. Production deployments, including work with Singapore Airlines, showed that internal tests often missed domain terminology, multi-part requests, and customer behavior patterns. Voice testing introduced further challenges such as accents, noise, interruptions, emotional states, agent handoffs, retrieval, latency, recovery, and human escalation. Salesforce simulates high call volumes with virtual calls and evaluates selected conversations for accuracy, completeness, conciseness, tone, and voice quality. Liang also identified retrieval among similar knowledge documents as a persistent weakness, arguing that continuous, representative, multi-turn testing in CI/CD is needed to measure and improve AI-agent performance across complete customer journeys.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.