Together AI brings Thinking Machines Lab’s new model Inkling on day 0
Blog post from Together AI
Inkling, a multimodal mixture-of-experts model developed by Thinking Machines Lab, is designed for token-efficient reasoning, native multimodal understanding, and versatility across a broad range of tasks, including scientific reasoning, coding, and calibrated prediction. Released on Together AI's inference platform, Inkling supports text, image, and audio inputs with text outputs through a unified decoder architecture that allows developers to control the inference effort to balance reasoning depth, token usage, and latency. It features architectural innovations such as query-conditioned relative attention, short causal convolutions, and a shared expert sink within a mixture-of-experts framework, enhancing its reasoning and multimodal capabilities. The model demonstrates strong performance in preliminary evaluations, excelling across various benchmarks including scientific reasoning, mathematical problem solving, and visual document understanding. Inkling's architecture incorporates a new approach to attention mechanisms and supports scalable, efficient execution, making it suitable for diverse applications like visual question answering and multimodal agents. It is available on Together AI Serverless, providing developers with immediate access and the ability to scale from experimentation to production without the need for extensive infrastructure management.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 7 | 1,111 | 224 | 91 | -41% |
| Serverless | 5 | 345 | 112 | 59 | -66% |
| LLM | 1 | 3,751 | 612 | 168 | -39% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.