Configuring Dedicated Model Inference
Blog post from Together AI
Together AI's Dedicated Model Inference platform integrates endpoints, deployments, and configs, which are tied together by a capacity-aware traffic split to facilitate model operations such as rollouts, A/B tests, and shadow experiments without downtime. An endpoint has a stable identity that applications use to call models, while deployments link specific models to configs and manage replicas with autoscaling policies. Configs specify the model's operational parameters, including engine type, GPU usage, and optimization profiles, and are immutable, ensuring consistent deployment behavior. Traffic routing is managed through weights assigned to deployments, allowing for proportional traffic flow based on capacity, which adapts to changes in replica counts. The platform allows for various traffic management strategies such as A/B testing and canary deployments, with traffic splits, A/B member percentages, and canary step percentages serving distinct roles in the routing process. The Together AI platform automates deployment setup and traffic routing, ensuring that new deployments do not receive traffic until specified in the endpoint's traffic split. Users can choose from certified profiles for their models, optimizing for latency, throughput, or a balance of both, depending on the use case, and can conduct tests to determine the best configuration for their traffic needs.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.