Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Canary rollouts: upgrade models in production without downtime

Blog post from Together AI

Post Details
Company
Date Published
Author
Together AI
Word Count
4,092
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

Model rollouts provide a managed way to migrate live endpoint traffic from an existing deployment to a new model through canary, blue-green, or rolling strategies, replacing manual cutovers with staged capacity scaling, health checks, routing-propagation waits, and optional metric gates. Canary rollouts move traffic through defined percentages and can compare router-side latency, error rate, or in-flight requests between source and target deployments, automatically pausing for human review when regressions, unavailable metrics, capacity shortages, or policy constraints occur. Operators can create, start, monitor, pause, resume, promote, cancel, or reverse rollouts through the CLI, REST API, or Python SDK; cancellation freezes the current traffic split, while a reverse rollout moves traffic back safely. The platform is designed to prevent traffic from reaching unready capacity, preserve sufficient replicas for each traffic share, respect autoscaling policies, and avoid declaring a regressed step successful. In a demonstrated Qwen2.5-7B to Qwen3.5-9B migration, a p95 latency gate detected a 137% regression at 10% traffic, paused the rollout, and enabled reversal to the original model without failed requests across 6,800 live requests.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.