Home / Companies / Deepinfra / Blog / August 2026

August 2026 Summaries

12 posts from Deepinfra

Filter
Month: Year:
Post Summaries Back to Blog
DeepInfra has introduced Sandboxes, an on-demand code-execution service that provides isolated Linux microVMs for AI agents, data pipelines, evaluation systems, and user-facing code-running features. Using a Python SDK, developers can create a sandbox, execute shell commands or Python, transfer files through a persistent `/workspace` directory, and either stop, restart, or permanently terminate the environment. The service uses Kata Containers on QEMU/KVM to provide hardware-enforced isolation rather than shared-container separation, and offers synchronous and asynchronous APIs, typed automation errors, tags, sandbox reattachment by ID, and a limit of five active sandboxes per account. Sandboxes are billed by the second while running, with nano instances starting at $0.054 per hour, while stopped instances incur no compute charge; they automatically stop after configured idle timeouts or at most 24 hours, and retained workspaces can be resumed for up to seven days. Planned additions include streamed command output, snapshots, port exposure, and directory uploads.
Aug 19, 2026 1,304 words in the original blog post.
DeepInfra’s guide compares SaaS tools and API platforms for accessing and deploying Kimi K3, a 2.8-trillion-parameter multimodal model with a one-million-token context window, emphasizing trade-offs in inference cost, latency, infrastructure control, customization, reliability, and compliance. It presents DeepInfra as its overall recommendation based on optimized bare-metal inference, OpenAI-compatible APIs, private endpoints, support for JSON mode, function calling and multimodal input, and listed pricing that discounts cached prompts. Moonshot AI is positioned as the official provider with Day-0 updates, privacy features, and automatic context caching, while Puter targets frontend developers through backend-free JavaScript integration. AIMLAPI and CometAPI are described as multi-model aggregators, with CometAPI emphasizing routing and failover, whereas Modal, Together AI, Fireworks AI, and Baseten target engineering and research teams needing serverless GPUs, fine-tuning, high-throughput serving, dedicated hardware, or regulated-industry compliance. The guide concludes that platform selection should depend on a team’s technical expertise, budget, desired level of deployment control, and production requirements.
Aug 13, 2026 1,871 words in the original blog post.
Kimi K3, released by Moonshot AI in July 2026, is an open-weight 2.8-trillion-parameter mixture-of-experts model with 104 billion active parameters, multimodal input support, text output, and a 1,048,576-token context window, positioning it for long-context reasoning, coding, and agentic workloads. The DeepInfra analysis describes strong benchmark performance in reasoning, coding, and agent use, while noting that the model is comparatively expensive, slower, and prone to verbose outputs. Provider prices cluster around $2.80–$3.00 per million input tokens and $14.00–$15.00 per million output tokens, with DeepInfra offering $2.85 input, $14.25 output, and a substantially discounted $0.285 cached-input rate, OpenRouter presenting the lowest listed base price, and Kimi’s first-party API serving as a direct managed-access reference. The discussion emphasizes that output length, repeated context, caching, routing, and features such as function calling, JSON mode, multimodal support, and private endpoints can affect real costs more than small differences in headline rates. It presents DeepInfra as particularly suitable for recurring large-context tasks such as repository-scale coding agents, document-heavy retrieval systems, multimodal debugging, structured tool workflows, and evaluation pipelines, while suggesting OpenRouter for lower-cost short-lived calls and Kimi’s API for teams prioritizing direct access.
Aug 12, 2026 3,365 words in the original blog post.
DeepInfra announces availability of Kimi K3, Moonshot AI’s 2.8-trillion-parameter open-weight native multimodal Mixture-of-Experts model, designed for software engineering, agentic workflows, scientific research, and text, image, and video processing. Although it contains 896 experts, it activates 16 per token for 104 billion active parameters, with features including Kimi Delta Attention, attention residuals, quantization-aware training, and a one-million-token context window intended to improve long-context efficiency. DeepInfra reports that benchmark results place Kimi K3 competitively against leading proprietary models in reasoning, coding, browsing, and visual tasks, while highlighting its tool-calling reliability, document analysis, and video capabilities. Developers can access the model through DeepInfra’s OpenAI-compatible chat-completions API using text or image inputs, JSON output, function calling, and configurable generation parameters. Pricing is usage-based at $2.85 per million input tokens, $14.25 per million output tokens, and $0.285 per million cached input tokens, positioning the service as a lower-cost option for repeated long-context workloads.
Aug 11, 2026 1,220 words in the original blog post.
DeepInfra has launched day-zero serverless API access to NVIDIA Nemotron 3.5 Lightning, an open 30-billion-parameter hybrid Mixture-of-Experts model designed for high-volume, always-on AI agents. Activating 3 billion parameters per token, the model supports up to 1 million tokens of context and uses DFlash speculative decoding and multi-token prediction to accelerate structured outputs such as tool calls and multi-step plans. NVIDIA and DeepInfra claim it delivers up to four times the throughput of comparable open models and up to 30% faster agent task completion, with reported benchmark scores of 86.5% on PinchBench agent productivity and 69.9% on AA-Omniscience Non-Hallucination. It is positioned as a customizable, high-throughput model for specialized workflows in areas including personal assistance, finance, cybersecurity, telecom, and retail, potentially alongside model-routing systems such as NVIDIA NeMo Switchyard. The OpenAI-compatible endpoint supports streaming, tool calling, and JSON mode, costs $0.05 per million input tokens and $0.20 per million output tokens, and DeepInfra states that it does not retain requests or use customer data for model training.
Aug 11, 2026 1,147 words in the original blog post.
DeepInfra has introduced Moonshot AI’s Kimi K3, an open-weight sparse Mixture-of-Experts model with 2.8 trillion total parameters but 104 billion activated per token, designed for long-horizon coding, agentic work, multimodal reasoning, and large-context applications. The model uses Kimi Delta Attention, Attention Residuals, and a Stable LatentMoE design that routes tokens across 16 of 896 experts, which Moonshot reports improves scaling efficiency over Kimi K2, while providing a native context window of roughly one million tokens and image and video support through the MoonViT-V2 encoder. DeepInfra reports strong benchmark results, including a leading 42.0 score on SWE-Marathon, while noting that K3 uses thinking mode by default, requires full assistant-message history in multi-turn tool workflows, and may need explicit constraints for autonomous agent tasks. Available through DeepInfra’s OpenAI-compatible API as moonshotai/Kimi-K3, it supports JSON output, function calling, and multimodal inputs, with usage-based token pricing, cache discounts, optional private deployments, and the provider’s stated zero-retention and security certifications.
Aug 10, 2026 1,346 words in the original blog post.
Data sovereignty has become a central requirement for AI deployments in regulated sectors because organizations must determine not only where data is stored, but also which jurisdiction can compel access to prompts, model inputs, logs, embeddings, and other derived data. The discussion distinguishes data residency, localization, and sovereignty, arguing that selecting a geographic region alone may not resolve legal exposure when infrastructure is operated by a company subject to another country’s laws. It presents open-weight models as a way for organizations to move models into cloud accounts, private endpoints, or on-premises environments they control, while noting that open weights and open-source licensing are not identical and may carry commercial restrictions. The deployment options range from shared APIs with zero-retention policies to dedicated endpoints, self-hosting in a customer cloud, and air-gapped infrastructure, with increasing operational responsibility at each tier. It also emphasizes that compliance reviews should inventory overlooked data surfaces such as retrieval chunks, embeddings, caches, observability traces, and fine-tuning datasets, since provider retention policies cannot protect information copied into an organization’s own logs. DeepInfra argues that OpenAI-compatible APIs can allow teams to change hosting boundaries with minimal application changes, and frames cost and model quality as secondary considerations after a deployment meets sovereignty and compliance obligations.
Aug 07, 2026 2,447 words in the original blog post.
DeepInfra’s guide distinguishes prompting, retrieval-augmented generation (RAG), fine-tuning, and model distillation as complementary approaches for improving production AI systems according to their specific failure modes. Prompting is recommended first for clarifying instructions, constraining outputs, and establishing a baseline, while RAG is appropriate when models need access to current, private, permission-controlled, or source-grounded information at inference time. Fine-tuning is intended for stable, repeatable behavioral shortcomings such as unreliable formatting, extraction, tool use, or domain-specific transformations, with LoRA presented as a lower-cost practical alternative to full fine-tuning. RAG and fine-tuning can be combined when systems need both changing external knowledge and consistent specialized behavior, but retrieval and generation should be evaluated separately to diagnose errors accurately. Distillation should follow only after a workflow is validated and stable, using a smaller model to reduce latency, memory use, and cost while testing quality against unseen and edge-case production data. DeepInfra positions its OpenAI-compatible platform as supporting this progression through model hosting, prompt caching, embeddings, rerankers, LoRA deployments, private models, and inference for distilled models.
Aug 06, 2026 2,942 words in the original blog post.
DeepInfra has introduced Prompt Cache Retention, a feature for Chat Completions and Text Completions that allows customers to explicitly keep reusable prompt context cached for either five minutes or one hour. Designed for agent workflows, multi-turn chats, and document question-answering, it can reduce time to first token and input costs by avoiding repeated prompt prefilling. Users enable retention with a stable prompt cache key and explicit TTL option, while cache breakpoints can restrict retention to a stable prompt prefix and exclude variable user questions. Initial cache writes carry premiums of 1.25 times standard input pricing for five minutes or 2.0 times for one hour, while later reuse is charged at the model’s discounted cache-read rate; cache windows can be extended but not shortened. The feature is currently supported on NVIDIA Nemotron-3-Ultra-550B-A55B and Moonshot AI Kimi-K2.7-Code, with response usage fields reporting cached and retained token counts.
Aug 05, 2026 1,376 words in the original blog post.
Open-source multimodal AI models should be evaluated beyond benchmark scores because production workloads involve messy documents and images, long contexts, multi-step tool use, latency under load, and token-driven costs that curated tests often miss. DeepInfra recommends Qwen3-VL for document extraction and multilingual OCR, highlighting its long-context handling and ability to interpret complex layouts; Kimi K3 for visual agents, GUI automation, and extended multi-step workflows; Gemma 4 26B for low-cost, high-volume image understanding and visual question answering; and MiMo-V2.5 for unified text, image, video, and audio pipelines. While these models have advanced rapidly, reliable noisy-audio transcription and long-form video reasoning remain challenging and may require specialized systems or extensive testing. Effective deployment also depends on API integration, balancing latency with throughput, managing context consumption, caching repeated inputs, and choosing quantization settings, with model selection and infrastructure decisions jointly determining real-world reliability and cost.
Aug 05, 2026 2,379 words in the original blog post.
DeepInfra's blog post discusses the comparison between vLLM and SGLang, two inference engines used for AI workloads, emphasizing that benchmark numbers often fail to provide a clear decision-making guide due to differences in configurations and workloads. The article highlights that choosing between vLLM and SGLang should be based on specific workload characteristics, such as prefix reuse, batch shape, structured output share, and model topology, rather than relying solely on throughput benchmarks. It also points out the importance of considering whether to self-host or use a managed endpoint, factoring in costs, regulatory requirements, workload types, and operational complexities. The post underscores the need to measure individual workload characteristics to make an informed decision and suggests that the choice between self-hosting and using a managed service depends on utilization patterns, regulatory needs, and operational priorities.
Aug 04, 2026 2,539 words in the original blog post.
DeepInfra's comparison between the GLM 5.2 and Claude Opus 4.8 models highlights a nuanced decision-making process for users prioritizing cost and performance in AI-driven tasks. While Claude Opus 4.8 demonstrates superior performance in coding benchmarks, particularly in long-horizon tasks, GLM 5.2 offers a more cost-effective solution due to its significantly lower price per task, making it suitable for bounded and verifiable work. The analysis reveals that the choice between these models should not be binary but rather task-dependent, with GLM 5.2 being favored for tasks where retries are feasible, and Claude Opus 4.8 for complex tasks requiring higher precision and fewer attempts. Additionally, the models differ in tokenization efficiency, with Claude Opus consuming more tokens due to a finer-grained approach, impacting the overall cost. DeepInfra suggests a strategic approach of using GLM 5.2 for general tasks and escalating to Claude Opus 4.8 for high-stakes tasks, emphasizing the importance of a dynamic routing strategy to maximize efficiency and cost-effectiveness.
Aug 03, 2026 2,318 words in the original blog post.