ClickHouse cluster operations in production
Blog post from Tinybird
ClickHouse cluster operations encompass adding or removing replicas, routing read and write workloads, resizing hardware, upgrading versions, managing replication and merges, and handling shards, with the central operational question being who can make changes and who bears responsibility when they fail. In self-hosted deployments, teams must manage Keeper, replica synchronization, load balancing, metadata cleanup, and risks such as replication lag, excessive data parts, and the lack of built-in online resharding; replicas duplicate data while shards divide it. Conventional managed offerings reduce infrastructure administration but still require customers to manage application behavior, data modeling, batching, and capacity policies. Tinybird’s dedicated Cluster Management feature is presented as a model in which customers choose replica counts, sizes, and traffic weights while Tinybird provisions, replicates, and operates the underlying cluster, with UI and API controls for adding, draining, removing, and routing traffic across replicas. The feature uses separate read, write, and copy-job weights with safeguards requiring at least one active read and write destination, but it is not autoscaling, on-demand sharding, or a replacement for query and schema optimization.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 1 | 3,490 | 385 | 112 | +26% |
| Observability | 1 | 3,175 | 737 | 186 | -24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.