Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Introducing the Flex Service Tier: Cheaper Inference When You Can Wait

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,099
Company Posts That Month
24
Language
English
Hacker News Points
-
Post removed?
No
Summary

DeepInfra has introduced a new Flex service tier designed to offer users a cost-effective alternative for non-urgent computational tasks, providing a 20% discount compared to real-time pricing. This tier is intended for workloads that prioritize cost over immediate response, such as model evaluations, data enrichment, and asynchronous tasks, by allowing a wait time of up to 10 minutes for available capacity. The Flex service operates through the same API as standard requests, making it easy for users to switch by simply adding "service_tier": "flex" to their requests. If the request waits without being processed, no charges are incurred, and this new feature is compatible with any OpenAI-compatible models that support the tier, offering a seamless transition for existing users.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 5 5,522 1,291 230 -4%
Vector Search 1 1,957 402 133 +3%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.