Home / Companies / DigitalOcean / Blog / Post Details
Content Deep Dive

How DigitalOcean’s Agentic Inference Cloud powered by NVIDIA GPUs Achieved 67% Lower Inference Costs for Workato

Blog post from DigitalOcean

Post Details
Company
Date Published
Author
Rithish Ramesh
Word Count
2,756
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

DigitalOcean's collaboration with Workato's AI Research Lab resulted in a significant reduction in inference costs and improved performance for Workato's automation processes using agentic AI capabilities. By deploying NVIDIA Dynamo with vLLM on DigitalOcean Kubernetes Service (DOKS), the team achieved a 67% lower inference cost by utilizing NVIDIA H200 GPUs, which provided enhanced memory capacity and efficient throughput. The key innovation was the implementation of KV-aware routing, which minimized redundant computations by leveraging warm KV caches, dramatically reducing latency and increasing throughput. This approach facilitated a 67% increase in tokens per second per GPU and reduced the number of GPUs needed by 40%, leading to substantial cost savings. The success of this project underscores the importance of optimizing the system architecture around AI models for efficient inference at scale, rather than merely relying on additional hardware.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 7 6,078 960 218 +18%
Kubernetes 4 1,840 308 106 +33%
Real-time 3 6,457 1,307 242 +28%
AI Agents 1 4,545 963 231 +27%
OpenClaw 1 650 79 49 -45%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.