Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Blog post from Together AI

Post Details
Company
Date Published
Author
Together AI
Word Count
1,008
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

Inkling, a multimodal mixture-of-experts model developed by Thinking Machines Lab, is designed for token-efficient reasoning, native multimodal understanding, and versatility across a broad range of tasks, including scientific reasoning, coding, and calibrated prediction. Released on Together AI's inference platform, Inkling supports text, image, and audio inputs with text outputs through a unified decoder architecture that allows developers to control the inference effort to balance reasoning depth, token usage, and latency. It features architectural innovations such as query-conditioned relative attention, short causal convolutions, and a shared expert sink within a mixture-of-experts framework, enhancing its reasoning and multimodal capabilities. The model demonstrates strong performance in preliminary evaluations, excelling across various benchmarks including scientific reasoning, mathematical problem solving, and visual document understanding. Inkling's architecture incorporates a new approach to attention mechanisms and supports scalable, efficient execution, making it suitable for diverse applications like visual question answering and multimodal agents. It is available on Together AI Serverless, providing developers with immediate access and the ability to scale from experimentation to production without the need for extensive infrastructure management.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 7 1,111 224 91 -41%
Serverless 5 345 112 59 -66%
LLM 1 3,751 612 168 -39%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.