Home / Companies / Coval / Blog / May 2026

May 2026 Summaries

12 posts from Coval

Filter
Month: Year:
Post Summaries Back to Blog
In 2026, Vapi significantly evolved its platform, securing a $50M Series B funding at a $500M valuation, and notably being chosen by Amazon Ring over 40 competitors for its voice AI capabilities. The company introduced a suite of new features, including Composer, a natural-language agent builder that expands Vapi's usability beyond engineers, and advanced monitoring tools for production observability. Vapi's offerings include Assistants and Squads for voice agent orchestration, while de-emphasizing Workflows. Integrations with platforms like GPT-5 and OpenAI Realtime highlight its commitment to maintaining an adaptable, open architecture. Vapi's pricing strategy remains usage-based, with native observability tools that provide detailed monitoring but are primarily beneficial for those committed to the Vapi ecosystem. The platform targets B2B-grade voice agents, offering flexibility and compliance options, though it faces competition from other platforms for specific low-latency or no-code needs.
May 28, 2026 3,141 words in the original blog post.
By 2026, ElevenLabs has evolved from a company known for its "shockingly human" text-to-speech (TTS) voices in 2023 to a $500 million annual recurring revenue audio AI platform with a diverse range of offerings, including voice cloning, speech recognition in over 90 languages, and a conversational agents platform. The company has introduced Instant Voice Cloning for quick applications and Professional Voice Cloning for high-quality, consistent output. It has also launched Eleven v3, its most expressive TTS model, alongside Flash v2.5, designed for real-time applications. The speech-to-text (STT) capabilities are covered by Scribe v2 and its Realtime variant, both offering fast, multi-language support. ElevenAgents, the rebranded production agent platform, includes features like Expressive Mode and Speech Engine, with recent pricing adjustments making the services more affordable. The company's expansion into enterprise and government sectors is supported by compliance with various regulatory frameworks and the introduction of AIUC-1 certification, which facilitates insurability for AI voice agents. ElevenLabs' footprint is marked by significant partnerships and deployments across large enterprises and government entities, with a comprehensive evaluation infrastructure enabling teams to assess its suitability for production environments.
May 25, 2026 3,143 words in the original blog post.
Arize and Coval together offer a comprehensive solution for evaluating voice AI applications by combining system-level observability with conversation-level simulation and evaluation capabilities. Arize captures detailed technical traces from voice AI systems, including internal system calls, audio processing events, and performance metrics, allowing for in-depth troubleshooting and performance monitoring. Coval utilizes these traces to provide advanced conversation-level analysis, simulation, and evaluation, facilitating the assessment of user experience and conversation quality. By setting up Arize for detailed system tracing and enabling Coval to pull and analyze this data, users can optimize both the technical aspects and conversational elements of their voice AI applications. This integrated approach ensures that users can maintain technical performance while simultaneously enhancing conversation quality, ultimately leading to improved user satisfaction.
May 22, 2026 1,178 words in the original blog post.
Continuous Integration and Continuous Delivery (CI/CD) for voice AI applies the same software engineering principles used in traditional software development to the development of voice agents, ensuring that every prompt, tool, and configuration change is version-controlled, tested, and deployed through automated gates. This approach helps prevent regressions and ensures compliance by treating voice agents as production software rather than mere configurations. Teams that implement CI/CD for voice AI benefit from reduced incident frequency, increased shipping velocity, and improved engineering quality of life, as they can confidently experiment and iterate on their voice agents. Key stages in the CI/CD pipeline include pre-commit checks, pull request evaluations, staging deployments, production rollouts with canary testing, and post-deployment monitoring, all supported by robust version control and evaluation platforms. By adopting a disciplined CI/CD methodology, voice AI teams can maintain high standards of quality and reliability, enabling them to scale their operations and address the unique challenges of voice AI development effectively.
May 21, 2026 3,156 words in the original blog post.
The text discusses the challenges and considerations involved in deciding whether to build or buy testing infrastructure for voice AI applications. Initially, many voice AI teams opt to create in-house testing scripts, which seem cost-effective in the short term but often become burdensome and inefficient over time, especially as complexity increases with factors like audio realism, multi-turn conversations, tool-call evaluations, scalability, regression discipline, and maintenance requirements. These challenges lead to predictable inflection points where internal solutions fall short, prompting teams to reconsider their approach. The text suggests that buying specialized voice AI evaluation platforms can be more cost-effective in the long run, freeing up engineering resources for more strategic work and offering better tools for scaling, concurrency, and realistic testing conditions. For most voice AI teams, the decision to buy rather than build is clearer when recognizing that testing infrastructure is typically a commodity rather than a differentiator. Nonetheless, there are scenarios, such as extreme volume or unique compliance needs, where building might still be the right choice. The guide emphasizes the importance of assessing whether testing infrastructure is a core differentiator for the business and provides a framework for making informed build-or-buy decisions.
May 19, 2026 3,301 words in the original blog post.
Voice AI regression testing is a critical practice for managing the dynamic and complex nature of voice agents, which are prone to silent regressions due to probabilistic outputs, multi-turn interactions, tool-call chains, model drift, and prompt sensitivity. This testing involves running a versioned library of conversational scenarios against a voice agent with every change, and comparing outcomes to a baseline to identify behavioral shifts. A robust regression suite should include a versioned scenario library, automated execution in CI/CD, behavioral grading, statistical thresholds, and diff views to effectively catch regressions and prevent the whack-a-mole cycle, where fixing one issue inadvertently causes another. The Coval platform offers a structured approach to voice AI regression testing, assisting teams in building an infrastructure that supports frequent, confident shipping by turning what could be a manual and error-prone process into an automated and reliable service. The methodology emphasizes the importance of realistic conversational scenarios, automation, and continuous integration of production data to enhance the testing suite's relevance and effectiveness over time.
May 16, 2026 3,328 words in the original blog post.
The text discusses the intricacies of evaluating text-to-speech (TTS) systems, emphasizing the significance of Time to First Audio (TTFA) and Word Error Rate (WER) as critical benchmarks for assessing performance. Coval, an independent evaluation platform founded by Brooke Hopkins, applies rigorous testing methodologies to provide standardized TTS benchmarks, enabling fair comparisons across providers. As of May 2026, Gradium leads these benchmarks, demonstrating superior latency and competitive WER, attributed to its innovative audio language models that integrate multiple voice tasks in a single architecture. The platform's open-source benchmarking tools allow developers to conduct their evaluations, offering a more realistic assessment of TTS performance in production environments compared to vendor-reported metrics.
May 14, 2026 1,088 words in the original blog post.
Voice AI systems often experience a significant performance gap between pre-launch testing and real-world production, with successful call handling dropping from 90-95% in testing to 60-70% in production. This discrepancy is primarily due to inadequate test coverage that fails to account for real-world conditions such as audio realism, accents, frustrated callers, integration surprises, and unknown variables. The text emphasizes the importance of a three-layer approach to bridge this gap, involving pre-launch simulation against realistic conditions, production observability to monitor ongoing performance, and a feedback loop that integrates production failures into new test scenarios. The Coval methodology, inspired by practices from the self-driving car industry, aims to enhance the accuracy of voice AI systems by continuously updating test scenarios based on real-world data, thus improving deployment speed and reducing incidents. The approach stresses the need for testing under realistic conditions, particularly focusing on audio realism, as the highest return on investment to ensure the agent's performance matches pre-launch expectations in varied and unpredictable production environments.
May 13, 2026 3,137 words in the original blog post.
Voice observability is a practice that assesses the functionality and quality of voice AI agents in production, focusing on both operational metrics like uptime and behavioral metrics such as conversation quality and customer satisfaction. Traditional observability tools like Datadog are adept at tracking operational metrics but fall short in evaluating the nuanced, multi-turn interactions of voice agents. Companies like Coval have emerged to fill this gap by providing a comprehensive observability stack that includes continuous grading of conversations, trace-level storage, real-time dashboards, and alerts, ensuring that issues are detected before they affect customers. The emphasis is on behavioral observability to maintain the quality of AI interactions, which involves assessing whether agents understood calls correctly, adhered to policy, and delivered satisfactory customer experiences. This approach not only helps in detecting failures that are invisible to operational metrics but also facilitates a feedback loop that integrates production insights into pre-production simulations, thus enhancing the overall reliability and performance of voice AI systems.
May 10, 2026 3,248 words in the original blog post.
AI agents are sophisticated systems that employ language models to autonomously execute multi-step tasks by utilizing external tools, distinguishing them from earlier AI systems such as chatbots and copilots through their autonomy, tool use, and goal-directed reasoning. These agents are increasingly deployed across various fields, including voice, chat, coding, and browser/research, each presenting unique evaluation challenges. Despite high success rates in controlled environments, there is often a significant drop in performance under real-world conditions, necessitating a robust evaluation framework based on functional correctness, tool use accuracy, behavioral quality, and safety/compliance. This continuous evaluation is crucial for identifying and mitigating failures that traditional QA processes might miss. The complexity of building an effective evaluation infrastructure often leads teams to consider purchasing rather than building these systems to better allocate engineering resources and ensure the agent's reliability and effectiveness in production environments.
May 08, 2026 3,757 words in the original blog post.
Voice AI models, encompassing components like speech-to-text (STT), large language models (LLM), and text-to-speech (TTS), are pivotal in creating effective voice agents, with significant considerations around latency, naturalness, and reliability. In 2026, the choice between cascaded pipelines and speech-to-speech models is critical; cascaded pipelines offer observability and separate evaluation of each component, while speech-to-speech models provide lower latency and more natural interaction by processing audio directly. The LLM layer varies significantly across deployments, with options like OpenAI's GPT-4o and GPT-Realtime-2, Anthropic's Claude, and Google's Gemini each offering distinct advantages in terms of speed, cost, and capability. Model selection should be driven by empirical testing on representative datasets to account for unique constraints like latency, cost, and multilingual capabilities, with a strong emphasis on real-world performance rather than benchmark scores. Teams are encouraged to build robust evaluation infrastructures to regularly test and optimize model choices, ensuring that the models meet the specific needs of their voice AI applications.
May 04, 2026 4,054 words in the original blog post.
In 2026, voice AI assistants have evolved significantly from their 2020 counterparts, transforming from simple voice retrieval systems to sophisticated agents capable of executing complex tasks such as scheduling, support, and transactions across various sectors including healthcare, banking, and customer service. Despite their potential, a gap often exists between demo performance and real-world deployment, with only 62% of voice AI systems surviving their first week in production due to unforeseen conditions and integration challenges. The two primary architectures for voice AI are cascaded pipelines and speech-to-speech models, each offering different advantages in terms of flexibility and latency. Successful deployment requires rigorous evaluation and testing strategies, including functional, behavioral, audio-environment, and adversarial testing to ensure reliability under diverse and challenging conditions. The market has consolidated around key platforms like Vapi, Retell AI, and LiveKit, which cater to different needs based on customization, vertical fit, and compliance requirements. Ultimately, the strategic focus for teams is on developing robust evaluation and monitoring frameworks to ensure voice AI assistants perform effectively in production, bridging the gap between technological capabilities and real-world application.
May 01, 2026 3,186 words in the original blog post.