Baseten leads Coval’s voice AI benchmark
Blog post from Baseten
Baseten reports that its deployment of Qwen3 ASR 1.7B Streaming leads Coval’s September 2026 early-access voice AI benchmark on the speech-to-text quality-latency Pareto frontier, achieving the lowest word error rate while operating about five times faster than OpenAI’s tested option. Coval provides reproducible, independent evaluations that account for datasets, model versions, normalization methods, and latency measurement, assessing STT through both transcription accuracy and speed and TTS through time to first audio. The report argues that voice-agent performance depends on the entire inference stack rather than model selection alone, as delays from STT, language models, TTS, networking, orchestration, voice activity detection, autoscaling, and provider-to-provider transfers compound during conversations. It emphasizes evaluating complete conversational turns and tail latency alongside median results, since inconsistent performance under load can degrade user experience. Baseten attributes reliable low-latency operation partly to optimized deployments, dedicated hardware, and infrastructure placement, while noting that benchmark results from a fixed region provide a baseline and real-world performance also depends on deployment architecture, location, workload type, quality requirements, and cost targets.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 12 | 324 | 41 | 16 | -89% |
| Real-time | 5 | 649 | 155 | 80 | -85% |
| LLM | 2 | 747 | 162 | 79 | -85% |
| AI Agents | 1 | 931 | 231 | 103 | -84% |
| AI Guardrails | 1 | 35 | 22 | 12 | -94% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.