How Lindy Cut Its Agent Inference Costs by 90% on Atlas Cloud
Blog post from Atlas Cloud
Lindy, a company developing AI “employees” that handle tasks across workplace tools such as Slack, calendars, meetings, and company records, reports moving much of its agent traffic from proprietary models including Claude, Sonnet, and Gemini to the open-weight DeepSeek v4 Flash model on Atlas Cloud. According to Lindy, the migration reduced inference costs by about 90% while supporting more than 3,000 requests per minute and over tenfold traffic growth, aided by caching that served roughly 60% of input tokens at lower rates. The company says Atlas’s unified API allowed it to test and deploy models from multiple providers quickly, ultimately using 24 models from 11 labs, while its direct contract, SOC 2 certification, and stated data-handling policies addressed security concerns for agents processing business data. The account argues that dedicated infrastructure provides more consistent model quality, capacity, observability, and support than marketplace routers, where requests may be handled by varying providers and serving configurations.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.