Inkling by Thinking Machines now available on Modal
Blog post from Modal
Inkling, released by Thinking Machines, is a general-purpose multimodal model designed to handle text, image, and audio inputs while generating text outputs. It features a mixture-of-experts transformer architecture with 975 billion total parameters, of which 41 billion are active, and employs a 1 million token context window with native audio and vision capabilities, prioritizing breadth over depth. The model is optimized for speed and efficiency through sparse experts and a unique local attention layout, where five out of every six attention layers use sliding window attention, enhancing performance by focusing on recent tokens. This design is supported by DFlash speculation, a technique that advances speculative decoding by generating whole blocks of tokens in parallel, maintaining speed and computational efficiency. Inkling is available on Modal as a Managed Endpoint with token-based pricing, promising improved interactivity and throughput on agentic workloads, and continues to evolve with advancements in local attention and speculative decoding techniques.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.