Introducing preemptible compute: the same compute, half the price
Blog post from Together AI
Together has launched a public preview of preemptible compute for Kubernetes-based GPU Clusters in all regions, offering interruptible NVIDIA GPU capacity at a fixed 50% discount from on-demand pricing with sub-hourly billing. Preemptible nodes join existing clusters rather than requiring a separate cluster type, use unused capacity, and may be reclaimed with up to five minutes’ notice through Kubernetes events and SIGTERM signals, allowing workloads time to checkpoint and exit before removal. The service automatically replenishes the requested preemptible capacity when resources become available, although allocated capacity may remain below the target during shortages. It is intended for retryable or checkpointed workloads such as experiments, training sweeps, evaluations, batch inference, and temporary demand bursts, while critical services, coordinators, login pods, and strict-SLO serving should remain on standard non-preemptible nodes. Users can configure preemptible GPU targets through the Together Cloud console, CLI, or API and schedule workloads using a compute-class node label; clusters must retain at least one standard node, and Together plans future Slurm support, expanded regions, and in-place conversion between compute types.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 5 | 956 | 75 | 30 | -73% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.