Introducing Prompt Cache Retention: Keep Your Context Warm for 5 Minutes or an Hour
Blog post from Deepinfra
DeepInfra has introduced Prompt Cache Retention, a feature for Chat Completions and Text Completions that allows customers to explicitly keep reusable prompt context cached for either five minutes or one hour. Designed for agent workflows, multi-turn chats, and document question-answering, it can reduce time to first token and input costs by avoiding repeated prompt prefilling. Users enable retention with a stable prompt cache key and explicit TTL option, while cache breakpoints can restrict retention to a stable prompt prefix and exclude variable user questions. Initial cache writes carry premiums of 1.25 times standard input pricing for five minutes or 2.0 times for one hour, while later reuse is charged at the model’s discounted cache-read rate; cache windows can be extended but not shortened. The feature is currently supported on NVIDIA Nemotron-3-Ultra-550B-A55B and Moonshot AI Kimi-K2.7-Code, with response usage fields reporting cached and retained token counts.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 4 | 5,068 | 1,020 | 229 | -34% |
| Real-time | 2 | 4,432 | 1,050 | 222 | -31% |
| AI Agents | 1 | 5,780 | 1,243 | 245 | -15% |
| Multi-agent systems | 1 | 432 | 163 | 64 | -19% |
| RAG | 1 | 1,152 | 209 | 75 | -6% |
| Vector Search | 1 | 2,358 | 371 | 127 | +5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.