October 2026 Summaries
5 posts from Deepinfra
Filter
Month:
Year:
Post Summaries
Back to Blog
GLM-5.3 is an open-weights Mixture-of-Experts reasoning model from Z.ai with 753 billion total parameters, 40 billion active parameters per token, a context window of roughly one million tokens, and support for text generation, structured outputs, and function calling through hosted providers. Positioned as a capable but relatively expensive, slow, and verbose model, it is aimed at long-context coding, tool use, agent workflows, and document-heavy retrieval applications rather than low-cost general-purpose use. Pricing varies substantially by provider, with list rates cited around $1.40 per million input tokens and $4.40 per million output tokens, while DeepInfra advertises standard rates of $0.90 and $3.00 respectively and a lower-priced Flex tier at $0.72 and $2.40; cached-input pricing can further affect costs for repeated-context workloads. The discussion emphasizes that output volume, reasoning settings, caching, latency, and deployment features such as private endpoints, zero data retention, and OpenAI-compatible routing may matter as much as headline input prices. DeepInfra is presented as particularly competitive for managed high-volume or batch deployments, while OpenRouter is described as offering easier multi-provider access, failover, and potentially faster routes, and self-hosting remains an option for organizations able to manage the associated infrastructure and operational costs.
Oct 03, 2026
3,474 words in the original blog post.
DeepInfra has made Z.ai’s GLM-5.3 reasoning model available through its platform, positioning it for complex software engineering, cybersecurity, and long-horizon agentic workflows. Built on the same base model as GLM-5.2, GLM-5.3’s reported improvements come solely from expanded reinforcement-learning post-training, using long-context processing, long-horizon optimization, and asynchronous training infrastructure. Z.ai reports substantial gains in coding and security benchmarks, including stronger results on Terminal Bench, DeepSWE, CyberGym, and ExploitGym, though many figures originate from internal or company-run evaluations and are not independently reproducible. The model supports a roughly one-million-token context window, function calling, JSON output, and three always-on reasoning settings, with maximum effort recommended for coding tasks. Z.ai also reports that its models have identified thousands of vulnerabilities in real codebases, although most findings remain under embargo. On DeepInfra, GLM-5.3 is offered through an OpenAI-compatible API with promotional, usage-based token pricing, optional Flex discounts, fp4 quantization, and platform security features including zero data retention and SOC 2 and ISO 27001 certifications.
Oct 03, 2026
1,854 words in the original blog post.
DeepInfra’s guide compares API providers and related tooling for deploying DeepSeek-V4.1-Flash, a model presented as having a 1M-plus-token context window and Engram conditional memory, with selection driven by latency, throughput, context caching, cost, availability, and compliance requirements. It positions DeepInfra as a balanced option for low-latency, cost-conscious production deployments, while the official DeepSeek API is recommended for repeated large-context workloads because of claimed cache-hit discounts. Canopy Wave and Atlas Cloud are highlighted for privacy and regulated-use requirements, OpenRouter and EvoLink.AI for routing, fallbacks, and avoiding lock-in, Fireworks AI and Novita AI for high-speed agentic or generation-heavy applications, Together AI for globally distributed enterprise reliability, and Clarifai for managed deployment with prompt-testing tools. Apidog is described not as a hosting provider but as a testing platform for comparing model outputs, latency, and token use before production releases.
Oct 02, 2026
1,936 words in the original blog post.
DeepInfra’s review describes GLM-5.3, Z AI’s August 2026 open-weights flagship model for coding, agentic reasoning, and cybersecurity, as a 753-billion-parameter mixture-of-experts model with 40 billion active parameters, a one-million-token context window, and an always-on reasoning engine. It reports an Artificial Analysis Intelligence Index score of 45 versus a comparable median of 18, alongside major benchmark gains over GLM-5.2, including Terminal-Bench 3.0 and CyberGym, but notes that the model is verbose and has relatively modest native output speed of 63.4 tokens per second. The native Z AI API is presented as the baseline, with a 3.43-second time to first token and blended pricing of $0.90 per million tokens, while the comparison of 14 providers identifies DeepInfra as the lowest-cost option at $0.72 per million blended tokens, Inco (FAST) as the throughput leader at 409.5 tokens per second, Databricks as the lowest-latency third-party provider at 9.38 seconds, and Fireworks as a high-throughput alternative for large-context workloads. The review advises teams to choose providers according to cost, output speed, latency, context requirements, capacity guarantees, and differing API behavior, including the removal of support for disabling GLM-5.3’s reasoning process.
Oct 02, 2026
1,738 words in the original blog post.
DeepInfra’s overview introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with a 552-billion-parameter backbone, text and image inputs, text outputs, and a native context window of 1,048,576 tokens. Its Causal Encoder-Decoder architecture activates 8 billion parameters during prompt prefill and 16 billion during decoding, while Compressed Sparse Attention 2 and bounded replay are presented as reducing KV-cache memory use to 890 bytes per token. The model also offers adjustable reasoning effort from 1 to 100, allowing users to trade speed and cost against reasoning depth. Internal benchmark results indicate strengths in coding, mathematics, automation, and selected agent evaluations, though it trails competing models on some difficult terminal, general-question-answering, and factual benchmarks, and results can vary with the evaluation scaffold. DeepInfra serves the model through an OpenAI-compatible API with function calling, JSON output, streaming, image-message support, and standard, priority, and flex service tiers. Standard pricing is listed at $0.20 per million input tokens, $0.60 per million output tokens, and $0.006 per million cached input tokens, with a typical output cap of 16,384 tokens despite the much larger total context capacity.
Oct 01, 2026
1,574 words in the original blog post.