5 recommendations when running Thanos and Prometheus
Blog post from Zapier
We have built a robust monitoring system using Thanos and Prometheus that provides fast and reliable monitoring capabilities, with high availability and scalability. We have learned the importance of caching to improve query performance, downsampling metrics to reduce storage requirements, keeping metrics in good shape by only storing important ones, sharding long-term storage to serve large amounts of data efficiently, and scaling and high availability through manual scaling of Prometheus shards. By implementing these strategies, we have improved the performance and reliability of our monitoring system, which currently stores 130TB of metrics and serves a high volume of queries within a few seconds.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 1 | 1,398 | 143 | 60 | +21% |
| Observability | 1 | 1,049 | 196 | 65 | +41% |
| OpenTelemetry | 1 | 201 | 26 | 14 | -16% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.