EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios
Blog post from Hugging Face
EVA-Bench Data 2.0 is an expansive framework designed to evaluate voice agents across three enterprise domains: Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR Service Delivery, covering 213 scenarios with 121 tools. This release, which represents a fourfold increase in scenario coverage from its original version, aims to test the adaptability of voice agents to varying domain-specific challenges such as vocabulary, workflow complexity, and user expectations. EVA-Bench is validated for solvability against advanced language models like OpenAI GPT-5.4, Google Gemini 3.1 Pro, and Anthropic Claude Opus 4.6, ensuring a robust and fair evaluation process. The dataset is structured to reflect realistic call patterns, including single and multi-intent scenarios, and adversarial interactions, with an emphasis on authentication and reproducibility. Scenarios are generated and validated using a graph-based pipeline to ensure consistency, and the datasets are open-source with multilingual support to address the challenges of deploying voice agents in diverse linguistic environments.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.