How Upstash Monitors Every Redis Replica with Checkly
Blog post from Checkly
Ilter Kavlak, a Site Reliability Engineer at Upstash, discusses the comprehensive monitoring strategy implemented at Upstash to proactively detect database issues before customers do, using a system of Checkly checks. This external monitoring layer, which operates globally from 18 locations every minute, ensures each database replica is individually checked for uptime, with the monitoring configuration managed as code in Terraform. The setup includes hundreds of URL monitors, primarily checking the /ping endpoint of each replica to confirm its functionality. Alerts are differentiated based on the severity of the issue, with Opsgenie and Slack channels handling notifications for failures and degraded states. The monitoring framework extends beyond simple HTTPS checks; for products like QStash, which handles message delivery, the system performs end-to-end checks to ensure messages are successfully processed. Internal monitoring layers complement this by checking the health of the software and tracking latency, but the external checks remain the ultimate indicator of customer experience. The approach emphasizes redundancy verification, regional monitoring, treating slowness as a distinct signal, and relying on external validation to declare incidents resolved, providing swift and clear insights for rapid incident response.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.