Autoscaling with GPU Transcription models
Blog post from Speechmatics
The company Speechmatics has moved its transcription models to use GPUs, which significantly improves accuracy but also increases costs. To ensure efficient processing and maintain cost-effectiveness, they use the Real Time Factor (RTF) ratio as a guideline for performance. However, running on GPU hardware introduces challenges such as shared resources and unpredictability in traffic demand. To address this, Speechmatics uses Kubernetes Event-Drive Autoscaling (KEDA) with Prometheus integration to scale their GPUs based on metrics provided by the Triton Server's `/metrics` endpoint. KEDA allows them to scale out based on specific metrics, including inference queue duration and count, which provides a more accurate representation of performance issues. Additionally, they implement deallocation on scale-down mode in AKS to accelerate node scaling and reduce pending time, resulting in improved cost efficiency and reliability.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 10 | 1,682 | 185 | 78 | +20% |
| Real-time | 1 | 2,062 | 598 | 178 | +12% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.