June 2026 Summaries
6 posts from Coval
Filter
Month:
Year:
Post Summaries
Back to Blog
Coval envisions a future where voice agents become a necessity for businesses of all sizes, similar to having a website today, by emphasizing the importance of trust and continuous improvement. The company focuses on creating self-improving voice AI systems that undergo rigorous testing and evaluation to ensure quality and reliability, akin to the development process of autonomous vehicles. Coval's approach involves a recursive self-improvement loop, where agents are continuously optimized through defining behaviors, testing, fixing issues, and validating improvements. This system is designed to transition from manual quality assurance to automated processes at scale, with features like metrics optimization and real-time insights. Coval aims to provide agents that not only detect and propose fixes for issues but also implement and measure these improvements, ultimately leading to highly efficient and trustworthy voice agents. The company's broader goal is to spearhead advancements in voice AI by fostering a collaborative community focused on building self-healing agent infrastructure.
Jun 29, 2026
1,102 words in the original blog post.
Brooke Hopkins, founder and CEO of Coval, discusses the company's mission to enhance the reliability of voice AI agents in production, inspired by her past experience at Waymo where robust infrastructure was crucial for deploying self-driving cars safely. Coval recently secured $28 million in Series A funding led by Norwest, with support from Base10 Partners, Twilio Ventures, Y Combinator, MaC Ventures, and Swift Ventures, bringing their total funding to $31 million. As enterprises increasingly integrate voice AI for customer interactions, the industry faces challenges in maintaining consistent performance beyond initial demos, with many agents failing due to inadequate infrastructure rather than model quality. Coval aims to address this gap by focusing on infrastructure to ensure voice agents work reliably and compliantly. The funding will support expanding Coval's sales and engineering teams and enhancing product capabilities, including improved simulations and integrations to provide seamless deployment in existing enterprise systems, ultimately preparing companies for the widespread adoption of voice AI as a primary interface with AI technologies.
Jun 23, 2026
674 words in the original blog post.
Running a voice AI vendor bake-off involves evaluating different voice AI platforms under identical conditions using a structured test suite that simulates real customer interactions. The aim is to select a vendor that provides the best user experience, rather than one that excels in scripted demos. This process mitigates common pitfalls such as selecting based on subjective impressions, demo theater, or overlooking rare but critical conversation types. The bake-off uses a frequency-weighted scenario and persona matrix derived from actual production calls to ensure realistic testing. Coval is highlighted as a neutral evaluation platform that facilitates this process by automating the creation and execution of test suites, ensuring a fair and defensible comparison. The method involves defining evaluation criteria upfront, running vendors through identical test conditions, scoring them using consistent metrics, and deciding based on pre-determined weights. This rigorous approach ensures that the selected vendor performs reliably in real-world conditions, with the added benefit of using the test suite for ongoing production monitoring after deployment.
Jun 15, 2026
6,212 words in the original blog post.
In Coval, human review is utilized to calibrate LLM-as-a-judge metrics, ensuring accurate evaluation of AI agents by validating the AI's scoring with human oversight. By having a QA team label a small sample of conversations and identify discrepancies between human and AI assessments, the calibration loop is established: labeling conversations, checking agreement rates, inspecting disagreements, fixing metric prompts, and re-testing calls. This process helps uncover issues like ambiguous prompts or incorrect AI judgments, enabling teams to refine metrics without altering the AI model or agent behavior. For example, human review revealed that a supposed compliance issue in an outbound voice agent was due to an incomplete metric prompt rather than actual agent failure. By refining this prompt, the AI's judgments aligned more closely with human assessments, reducing the need for continuous human oversight as the system's reliability improved. Coval also offers a service for teams to outsource this calibration process, ensuring metrics accurately reflect agent performance and support informed decision-making.
Jun 11, 2026
1,180 words in the original blog post.
In 2026, the speech-to-text (STT) landscape reflects significant shifts from previous years, with Word Error Rates (WER) on clean English audio plateauing among top providers like Deepgram, AssemblyAI, and OpenAI, all within 1-2 percentage points of each other. The competitive focus has moved towards features such as streaming latency, end-of-turn detection, multilingual support, and cost efficiency. Microsoft launched its first proprietary STT model, MAI-Transcribe-1, claiming a 3.8% avg WER across 25 languages and significantly reduced GPU costs. OpenAI introduced GPT-Realtime-Whisper, marking its first separation of streaming-optimized STT from the batch Whisper line. Meanwhile, Deepgram's Flux Multilingual STT model integrates end-of-turn detection without external VAD, enhancing response times. The market sees a variety of offerings, with providers like NVIDIA, Gladia, and Cartesia focusing on multilingual breadth and cost-effective solutions. Vendor benchmarks often don't accurately predict real-world performance due to variable conditions such as audio quality and language complexity. The guide emphasizes the importance of independent benchmarking using real traffic to evaluate STT providers effectively, considering factors like streaming latency, cost, and entity preservation, and suggests employing a multi-provider strategy for optimal results.
Jun 04, 2026
4,029 words in the original blog post.
In 2026, the text-to-speech (TTS) landscape has evolved significantly, with a market broken into three tiers: expressive offline models, real-time agent models, and high-volume cheap models. Key players like ElevenLabs, Cartesia, and OpenAI have made significant advancements, including ElevenLabs releasing its expressive Eleven v3 model and Cartesia offering the low-latency Sonic-3. OpenAI has integrated its TTS and STT functions into a single model with GPT-5-class reasoning. Vendor-reported benchmarks are often unreliable for real-world applications, making independent measurement against actual traffic essential for accurate assessment. The industry sees a shift from latency as the main differentiator to factors like emotional control, prosody, multilingual fidelity, and cost. Multi-provider strategies have become the norm for large-scale operations, allowing for fallback options, continuous evaluation, and quality optimization. The market has also seen price reductions, such as ElevenLabs' pricing reset, and new offerings like Vapi Voices Beta, which provide cost-effective solutions for high-volume, low-stakes interactions.
Jun 01, 2026
3,719 words in the original blog post.