When Self-Hosted Voice AI Models Make Sense: Baseten x Coval
Blog post from Coval
A conversation between Baseten’s Tianshu Cheng and Coval CEO Brooke Hopkins examines when voice AI teams may outgrow managed APIs and consider dedicated or self-hosted inference. They identify rising costs, requirements for fine-tuned or proprietary models, and inconsistent tail latency as key triggers, emphasizing that P90–P99 delays can repeatedly affect callers during multi-turn conversations even when median latency appears acceptable. The discussion recommends a three-stage selection process using public benchmarks to narrow candidates, testing on representative internal data, and full-agent simulations before gradual production rollout, supported by continuously updated evaluation datasets that capture real failures while addressing privacy through redaction or synthetic examples. It also highlights end-to-end latency optimization through model co-location, smaller task-specific models, and emerging architectures that separate fast interaction from deeper background reasoning or use speech-to-speech systems. Both speakers argue that model performance should be evaluated against specific business workflows rather than broad benchmarks or word error rate alone, particularly for critical information such as names, numbers, and addresses, and they describe customization and multi-model deployment as growing priorities for production voice systems.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Voice AI | 52 | 324 | 41 | 16 | -89% |
| LLM | 17 | 747 | 162 | 79 | -85% |
| AI Model Fine-tuning | 10 | 139 | 28 | 14 | -75% |
| Local AI | 5 | 15 | 4 | 3 | -94% |
| Serverless | 4 | 156 | 54 | 28 | -80% |
| Real-time | 3 | 649 | 155 | 80 | -85% |
| Observability | 1 | 472 | 102 | 54 | -85% |
| Vector Search | 1 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.