Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Introducing preemptible compute: the same compute, half the price

Blog post from Together AI

Post Details
Company
Date Published
Author
Together AI
Word Count
856
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

Together has launched a public preview of preemptible compute for Kubernetes-based GPU Clusters in all regions, offering interruptible NVIDIA GPU capacity at a fixed 50% discount from on-demand pricing with sub-hourly billing. Preemptible nodes join existing clusters rather than requiring a separate cluster type, use unused capacity, and may be reclaimed with up to five minutes’ notice through Kubernetes events and SIGTERM signals, allowing workloads time to checkpoint and exit before removal. The service automatically replenishes the requested preemptible capacity when resources become available, although allocated capacity may remain below the target during shortages. It is intended for retryable or checkpointed workloads such as experiments, training sweeps, evaluations, batch inference, and temporary demand bursts, while critical services, coordinators, login pods, and strict-SLO serving should remain on standard non-preemptible nodes. Users can configure preemptible GPU targets through the Together Cloud console, CLI, or API and schedule workloads using a compute-class node label; clusters must retain at least one standard node, and Together plans future Slurm support, expanded regions, and in-place conversion between compute types.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 5 956 75 30 -73%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.