How to Evaluate Live & Voice Agents in ADK
Blog post from Google Cloud
Native live evaluation in Google’s Agent Development Kit (ADK) enables developers to test voice-based agents through simulated spoken conversations, helping assess multi-turn behavior, timing, tool use, context retention, and recovery from interaction issues before deployment. The workflow supports graph-based multi-agent systems, with session state and audio streams preserved across handoffs, and evaluation cases can use either goal-driven simulated personas or fixed, scripted user conversations. Developers configure live mode, audio synthesis, simulated-user behavior, voice, language, turn limits, and rubric-based scoring in a test configuration, allowing end-to-end trajectory quality and per-turn performance to be measured. Evaluations can run through the ADK CLI or programmatically in CI/CD pipelines, while ADK Web provides transcripts, audio playback, and interactive debugging tools to inspect both what an agent said and how it sounded.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 2 | 2,839 | 275 | 56 | -36% |
| LLM | 1 | 5,068 | 1,020 | 229 | -34% |
| Multi-agent systems | 1 | 432 | 163 | 64 | -19% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.