Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

Introducing Prompt Cache Retention: Keep Your Context Warm for 5 Minutes or an Hour

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,376
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

DeepInfra has introduced Prompt Cache Retention, a feature for Chat Completions and Text Completions that allows customers to explicitly keep reusable prompt context cached for either five minutes or one hour. Designed for agent workflows, multi-turn chats, and document question-answering, it can reduce time to first token and input costs by avoiding repeated prompt prefilling. Users enable retention with a stable prompt cache key and explicit TTL option, while cache breakpoints can restrict retention to a stable prompt prefix and exclude variable user questions. Initial cache writes carry premiums of 1.25 times standard input pricing for five minutes or 2.0 times for one hour, while later reuse is charged at the model’s discounted cache-read rate; cache windows can be extended but not shortened. The feature is currently supported on NVIDIA Nemotron-3-Ultra-550B-A55B and Moonshot AI Kimi-K2.7-Code, with response usage fields reporting cached and retained token counts.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 4 5,068 1,020 229 -34%
Real-time 2 4,432 1,050 222 -31%
AI Agents 1 5,780 1,243 245 -15%
Multi-agent systems 1 432 163 64 -19%
RAG 1 1,152 209 75 -6%
Vector Search 1 2,358 371 127 +5%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.