Home / Companies / Modal / Blog / Post Details
Content Deep Dive

Inkling by Thinking Machines now available on Modal

Blog post from Modal

Post Details
Company
Date Published
Author
-
Word Count
647
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Inkling, released by Thinking Machines, is a general-purpose multimodal model designed to handle text, image, and audio inputs while generating text outputs. It features a mixture-of-experts transformer architecture with 975 billion total parameters, of which 41 billion are active, and employs a 1 million token context window with native audio and vision capabilities, prioritizing breadth over depth. The model is optimized for speed and efficiency through sparse experts and a unique local attention layout, where five out of every six attention layers use sliding window attention, enhancing performance by focusing on recent tokens. This design is supported by DFlash speculation, a technique that advances speculative decoding by generating whole blocks of tokens in parallel, maintaining speed and computational efficiency. Inkling is available on Modal as a Managed Endpoint with token-based pricing, promising improved interactivity and throughput on agentic workloads, and continues to evolve with advancements in local attention and speculative decoding techniques.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.