Home / Companies / PagerDuty / Blog / November 2015

November 2015 Summaries

5 posts from PagerDuty

Filter
Month: Year:
Post Summaries Back to Blog
Monitoring systems have become crucial for digital businesses by detecting issues like slow API requests and network problems, but many IT operations teams still struggle with efficient incident response. Despite investing in various specialized monitoring tools, which 91% of operations teams use, there is a significant challenge in managing the overwhelming volume of alerts generated daily, leading to alert fatigue and missed critical incidents, as reported by 85% of teams. The manual and inefficient nature of current incident response processes, often relying on slow and unaccountable methods like email and Excel sheets, leaves 54% of IT teams dissatisfied. To address these challenges, teams are encouraged to plan their incident response, implement best practices, and leverage analytics to identify and strategize on pain points, enhancing the effectiveness of their monitoring systems in protecting uptime.
Nov 24, 2015 534 words in the original blog post.
PagerDuty encountered unexpected performance issues with their Apache Cassandra deployments across geographically distributed datacenters, specifically noticing read hotspots despite uniform data distribution and access patterns. Initially attributing the uneven load to hardware discrepancies, further investigation revealed that the issue stemmed from the system's tendency to select nodes with the fastest response times for read requests. This led to certain nodes, particularly those in the us-west-1 datacenter, being disproportionately burdened due to their proximity and response efficiency. The challenge was compounded by the inherent round-trip time (RTT) variances in their network topology, resulting in a skewed distribution where nodes A-C received 60% more reads compared to others. Despite considering potential solutions, such as redistributing reads or adjusting network distances, the most viable approach was to scale up the overburdened nodes to manage the load effectively. This experience highlighted the complexities of assumptions in system design and the universal nature of such phenomena in distributed systems that favor low-latency operations.
Nov 20, 2015 1,207 words in the original blog post.
Initially skeptical of "Women in Tech" events, the author unexpectedly found value in leading a Women's Leadership Circle at PagerDuty, highlighting the positive impact such gatherings can have. The event featured discussions and talks by women in tech leadership, which served to boost confidence, encourage self-improvement, and foster a supportive community. Key takeaways included the importance of not fearing failure, confidently marketing oneself, negotiating for what one deserves, and recognizing and changing behaviors that might undermine one's professional presence. The event emphasized the power of small commitments to bring about significant change, with participants encouraged to text their commitments to someone for accountability. While acknowledging that not all "Women in Tech" events are equally beneficial, the author encourages seeking out or creating groups that align with one's personal and professional goals.
Nov 10, 2015 734 words in the original blog post.
Businesses are increasingly adopting a DevOps model to foster a more democratic workplace and improve customer experience by enabling agile responses to real-time demands. This cultural shift emphasizes high-trust environments where development and operations teams collaborate transparently and fluidly, often using tools like shared source code and incident dashboards. Companies like Etsy exemplify the successful integration of these practices, promoting continuous learning through programs that encourage cross-team understanding and blameless post-mortems to analyze failures constructively. Such environments cultivate positive feedback loops, leading to faster deployment, better incident response, and empowered teams engaged with data-driven decision-making. This democratized approach encourages employees to take ownership of their roles and contribute to the company's direction, echoing democratic principles where everyone has a voice in shaping outcomes. Gene Kim highlights that leaders in DevOps environments foster routines that drive continuous improvement, ensuring that obstacles are systematically addressed. Tools such as PagerDuty facilitate this by providing centralized oversight and analytics, reinforcing a foundation of transparent communication and collective ownership.
Nov 05, 2015 801 words in the original blog post.
PagerDuty has announced the integration of Event Enrichment HQ, along with its Event Enrichment Platform (EEP), into its system to create a comprehensive event management and incident resolution platform. This collaboration aims to alleviate the challenges faced by operations teams due to excessive telemetry data from monitoring systems, which often leads to alert fatigue and missed alerts. By implementing EEP, teams can reduce noisy incidents by up to 94% through custom classification rules that filter non-urgent events, ensuring only critical alerts reach the appropriate on-call teams. Additionally, EEP enhances incident resolution by providing direct access to remediation steps in incident notifications, streamlining the process and reducing downtime. The integration is currently available as a closed beta, allowing interested parties to sign up for early access.
Nov 03, 2015 437 words in the original blog post.