Home / Companies / Deepinfra / Blog / June 2026

June 2026 Summaries

6 posts from Deepinfra

Filter
Month: Year:
Post Summaries Back to Blog
DeepInfra's strategic integration of NVIDIA's inference software stack, including components like TensorRT-LLM, Dynamo, and NVFP4, has significantly enhanced its operational efficiency, as evidenced by the successful deployment of DeepSeek V4 with a remarkable 4x performance improvement. By relying on NVIDIA's Blackwell-generation GPUs and optimizing their models through quantization, DeepInfra has achieved a substantial reduction in infrastructure costs while maintaining performance, allowing developers to benefit from ongoing improvements without additional effort. This approach underscores DeepInfra's commitment to leveraging cutting-edge technology to provide faster and more cost-effective solutions, making it a pioneering force in scalable model deployment.
Jun 30, 2026 714 words in the original blog post.
DeepInfra has introduced a new Priority Service Tier to enhance its inference cloud capabilities, allowing latency-critical traffic to move to the front of the queue during high-demand periods, ensuring faster processing for essential tasks. This tier, which costs 1.5 times the regular rate, is aimed at applications where immediate response is crucial, such as interactive user-facing apps and revenue-critical functions. The Priority Service is seamlessly integrated into the existing OpenAI-compatible API, requiring only a simple field addition to requests and ensuring that users are billed the premium rate only when the priority service is actually applied. Currently, this service is live for models on the vLLM stack, with plans to extend support to additional models, providing a clear indication on model pages whether they are Priority-enabled.
Jun 29, 2026 1,039 words in the original blog post.
DeepInfra has launched a Batch API that enables users to run large, non-urgent inference jobs at a 20% reduced cost compared to real-time pricing, making it suitable for tasks such as dataset evaluation, generating embeddings, and large-scale classification. The API is compatible with OpenAI's Batch API, allowing users to easily transition their existing workflows by uploading a JSONL file, creating a batch, and polling for completion. This approach is designed for use cases where immediate responses are unnecessary, offering a more cost-effective solution by trading off latency for throughput. The Batch API supports multiple endpoints, including completions and embeddings, and applies automatic discounts to batch requests. Users can start by pointing their OpenAI client at DeepInfra and following the detailed documentation for seamless integration.
Jun 19, 2026 1,051 words in the original blog post.
DeepInfra has announced the release of Step 3.7 Flash, a 198-billion-parameter sparse Mixture-of-Experts vision-language model optimized for agentic workflows, now available on their platform. This model is designed to execute complex tasks such as parsing financial reports and managing multi-step search loops with a focus on execution reliability over raw model quality. It supports a 256K context window and offers three reasoning levels to balance speed, cost, and depth per request. Step 3.7 Flash integrates a language backbone with a vision encoder for native image understanding, performing well in benchmarks like ClawEval-1.1 and SimpleVQA. The model is accessible through DeepInfra's OpenAI-compatible API, maintaining competitive pricing and ease of use for developers familiar with the platform.
Jun 12, 2026 910 words in the original blog post.
DeepInfra has announced the launch of NVIDIA Cosmos 3, an innovative open world foundation model for physical AI, available in two variants, Cosmos 3 Nano and Cosmos 3 Super, on its platform. Cosmos 3 distinguishes itself by employing reasoning before generating outputs, a crucial feature for developing safe and reliable physical AI systems such as robots and autonomous vehicles, and uses a Mixture-of-Transformer architecture that integrates an autoregressive reasoner with a diffusion-based generator. The model handles multimodal inputs and outputs, supporting tasks across text, image, video, audio, and action, which makes it suitable for synthetic data generation, policy training for robotics, and visual reasoning in smart city infrastructure. Both variants of Cosmos 3 are designed to support developers in building scalable and efficient AI systems and are available at competitive rates through DeepInfra's standard API, with no special configuration required for setup.
Jun 04, 2026 769 words in the original blog post.
DeepInfra has announced the launch of its new Nemotron 3 Ultra and Nemotron 3.5 Content Safety models, emphasizing a shift towards task completion speed and efficiency in agentic AI systems. The Nemotron 3 Ultra is designed for complex workflows, offering up to five times faster inference and 30% lower costs, while supporting up to one million tokens in context. It focuses on specialized roles like reasoning and orchestration instead of a one-size-fits-all model. The Nemotron 3.5 Content Safety model provides multimodal safety measures, handling text and images across 12 languages with minimal latency, serving as a guardrail for AI safety testing and evaluation. Both models are integrated into DeepInfra's existing API, allowing users to easily access these advancements without altering their current setup. The new models complement each other by combining efficient task processing with robust safety protocols, facilitating more reliable and cost-effective AI solutions on the DeepInfra platform.
Jun 04, 2026 827 words in the original blog post.