October 2026 Summaries
2 posts from Together AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Together AI has partnered with IBM and NVIDIA to deploy a dedicated large-scale AI inference cluster on IBM Cloud, using NVIDIA B300 GPUs and Spectrum-X Ethernet networking, with Together AI as its first customer. Under the arrangement, IBM supplies enterprise cloud infrastructure, NVIDIA provides specialized computing and networking hardware, and Together AI operates the inference platform for open AI models. The company says it already serves hundreds of trillions of tokens monthly to more than one million developers and is expanding capacity in anticipation of rapidly rising demand. The collaboration is intended to offer enterprises scalable, secure, reliable, and cost-efficient open-model inference while supporting data sovereignty and performance comparable to closed AI systems.
Oct 06, 2026
318 words in the original blog post.
Together Link is a tool that connects existing coding-agent environments, including Claude Code, Claude Desktop, Codex, OpenCode, and Pi, to Together AI’s open models with the aim of reducing model spending by more than 50% without changing established workflows. It positions open models such as Kimi K3 and GLM 5.3 for complex coding work, while faster models including GLM 5.3 Flash and DeepSeek V4.1 Flash handle routine tasks at lower cost. Installation requires a single command and uses a Together AI API key, while an Auto routing option selects a model once per session based on the initial task, preserving prompt caching and optionally incorporating Anthropic models when an Anthropic key is supplied. The service provides per-session cost comparisons against Opus 5.5, bills through existing Together AI pay-as-you-go or credit-pack accounts, and operates on Together’s serverless inference infrastructure.
Oct 05, 2026
462 words in the original blog post.