Introducing the Flex Service Tier: Cheaper Inference When You Can Wait
Blog post from Deepinfra
DeepInfra has introduced a new Flex service tier designed to offer users a cost-effective alternative for non-urgent computational tasks, providing a 20% discount compared to real-time pricing. This tier is intended for workloads that prioritize cost over immediate response, such as model evaluations, data enrichment, and asynchronous tasks, by allowing a wait time of up to 10 minutes for available capacity. The Flex service operates through the same API as standard requests, making it easy for users to switch by simply adding "service_tier": "flex" to their requests. If the request waits without being processed, no charges are incurred, and this new feature is compatible with any OpenAI-compatible models that support the tier, offering a seamless transition for existing users.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 5 | 5,522 | 1,291 | 230 | -4% |
| Vector Search | 1 | 1,957 | 402 | 133 | +3% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.