The production platform for open-weight AI inference
Blog post from Together AI
A new update to the inference platform offers comprehensive control over performance, cost, and quality, allowing users to deploy models quickly and efficiently without creating their own infrastructure. The platform supports multiple deployments behind a single endpoint, enabling safe updates with canary, blue-green, and rolling strategies, and testing on real traffic with A/B and shadow testing. It also introduces a closed beta for custom training, including reinforcement learning and fine-tuning, with seamless deployment to production. Open-weight models are emphasized for their control over performance, quality, and functionality, enabling enterprises to incorporate proprietary IP without exposure risks. The platform simplifies transitioning from experimentation to production, supporting open-weight, licensed closed-weight, and fine-tuned models, while offering flexibility in hardware, optimization profiles, and scalability. It features enhanced model caching for faster deployment and supports diverse autoscaling metrics. Users can maintain stable endpoints while evolving deployments, using A/B testing and shadow traffic to test changes safely. Additionally, the platform provides extensive observability tools and integrates custom training directly into the production workflow, aiming to streamline the entire model lifecycle.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.