July 2026 Summaries
2 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Moonshot has launched Kimi K3, a groundbreaking 2.8 trillion parameter multimodal model designed for extensive token context windows and native vision capabilities. With strategic partnerships with Modal and vLLM, K3 is available on day one with token-based pricing and an Auto Endpoint option for dedicated capacity. Kimi K3 employs a mixture-of-experts transformer architecture, enabling efficient scaling with active experts per token and a 1 million token context window. Its innovative features, such as Kimi Delta Attention and Attention Residuals, provide 2.5 times the scaling efficiency of its predecessor, K2. Moonshot's efforts in quantization-aware training and expert parallelism optimization ensure the model runs across diverse hardware platforms. Modal enhances K3's performance with a custom DFlash speculator, significantly speeding up decode time and boosting throughput on agentic tasks. Available through Modal's platform with a $30 monthly compute credit, Kimi K3 is poised to redefine open AI model capabilities with its impressive performance metrics and accessibility.
Jul 27, 2026
558 words in the original blog post.
Inkling, released by Thinking Machines, is a general-purpose multimodal model designed to handle text, image, and audio inputs while generating text outputs. It features a mixture-of-experts transformer architecture with 975 billion total parameters, of which 41 billion are active, and employs a 1 million token context window with native audio and vision capabilities, prioritizing breadth over depth. The model is optimized for speed and efficiency through sparse experts and a unique local attention layout, where five out of every six attention layers use sliding window attention, enhancing performance by focusing on recent tokens. This design is supported by DFlash speculation, a technique that advances speculative decoding by generating whole blocks of tokens in parallel, maintaining speed and computational efficiency. Inkling is available on Modal as a Managed Endpoint with token-based pricing, promising improved interactivity and throughput on agentic workloads, and continues to evolve with advancements in local attention and speculative decoding techniques.
Jul 15, 2026
647 words in the original blog post.