Home / Companies / DigitalOcean / Blog / July 2026

July 2026 Summaries

5 posts from DigitalOcean

Filter
Month: Year:
Post Summaries Back to Blog
DigitalOcean's launch of the Kimi K3 model marks a significant milestone in AI deployment, showcasing the complexities of integrating a large-scale model with 2.78 trillion parameters into their Inference Engine. The deployment involved selecting high-performance hardware like NVIDIA HGX B300 and AMD Instinct MI350x GPUs, optimizing server configurations, and collaborating with teams like Moonshot AI and Inferact to ensure robust performance and verification against industry benchmarks. The model's serving recipe was fine-tuned to maximize throughput and minimize latency while maintaining user experience, involving precise hardware tuning and memory management strategies. The open-source Kimi Vendor Verifier (KVV) project was crucial in ensuring accurate model serving across different vendors, highlighting the importance of correct implementation to maintain benchmark integrity. Special features like dynamic tool calls required unique integration efforts, ensuring compatibility without affecting other models. The comprehensive efforts led to a successful deployment of Kimi K3, available through DigitalOcean’s Serverless Inference platform, offering users a glimpse into the potential of advanced AI models in real-world applications.
Jul 30, 2026 3,018 words in the original blog post.
DigitalOcean's Inference Engine introduces a server-side tool called model synthesis, designed to optimize the cost and quality trade-offs in AI model usage by orchestrating multiple models to work together. By employing a panel of models and a synthesizer to combine their outputs, users can achieve higher quality results at a lower cost compared to using a single model. Benchmarking on the DRACO deep-research tasks highlighted the effectiveness of an open-source model configuration (GLM 5.2 + Kimi K2.6), which outperformed the single model Fable 5 at half the cost. This approach is particularly beneficial for tasks requiring comprehensive and evidence-based answers, as it allows for flexible configuration tailored to specific needs, offering cost-efficient solutions without compromising on quality. The Inference Engine's model synthesis tool is available in Public Preview, enabling users to select from optimized presets or customize their model configurations, providing a scalable solution for diverse AI application needs.
Jul 23, 2026 1,931 words in the original blog post.
Effective August 1, 2026, DigitalOcean will update the pricing for select NVIDIA and AMD GPU droplets in response to increased demand for advanced GPU capacity, aiming to maintain competitive infrastructure pricing. Customers utilizing on-demand GPU services will be billed at the new rates from this date, with changes reflected in their September invoices, while those with 12-month reserved plans will see no change unless they renew at the adjusted rates. Customers are encouraged to contact DigitalOcean's sales team for assistance in finding the most cost-effective options for their workloads and to explore reserving capacity to lock in lower rates.
Jul 21, 2026 470 words in the original blog post.
DigitalOcean has introduced Managed Weaviate, now in public preview, providing a streamlined solution for running Weaviate in production with predictable pricing starting at $20 per month. Weaviate, an open-source AI-native vector database, is crucial for applications like semantic search and retrieval-augmented generation, supporting agent-driven workflows and similarity-based recommendations. Managed Weaviate simplifies operational tasks such as backups, security patching, and high availability, allowing developers to focus on building applications without managing infrastructure. It offers full API compatibility and integrates seamlessly into DigitalOcean's ecosystem, ensuring cost-efficiency and eliminating vendor lock-in. The service employs RQ8 compression to optimize storage and supports features like auto version upgrades and credential rotation. As Weaviate's adoption grows, DigitalOcean aims to enhance infrastructure support, making it increasingly effortless for developers to deploy AI-native applications.
Jul 09, 2026 917 words in the original blog post.
DigitalOcean's new Evaluations feature allows teams to validate models and inference routers on their own data before deploying them in production, ensuring optimal performance in terms of quality, latency, and cost. This service, part of the DigitalOcean Inference Engine, includes structured LLM-as-a-Judge evaluations, which score models against pre-built and custom metrics such as correctness, completeness, and bias. The platform supports managing datasets, creating reusable evaluation presets, and triggering evaluations programmatically, facilitating integration into continuous integration (CI) pipelines. Teams can access a range of judge models, including those from DigitalOcean's Model Catalog and external imports, with the option to upgrade for premium model access. Evaluations are integrated directly into the DigitalOcean stack, allowing validation against the same endpoints used in production, and are designed to be flexible, scalable, and repeatable as models evolve.
Jul 01, 2026 819 words in the original blog post.