September 2017 Summaries
9 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
A gathering in Silicon Valley on September 5 brought together leading innovators, CEOs, and founders to promote inclusivity and diverse thinking in the tech industry, featuring prominent women leaders sharing their insights and personal journeys. The event highlighted the challenges of creating an inclusive environment where everyone feels valued, with discussions led by notable figures such as Merline Saintil, Pamela Rice, Yvonne Wassenaar, Sheila Jordan, and Alvina Antar, who emphasized the importance of advocacy, self-belief, sponsorship, cultural shift, and organic diversity. The panelists shared personal anecdotes and advice on overcoming adversity, leveraging mentorship and sponsorship, and operationalizing diversity initiatives within organizations. The event underscored the necessity for both men and women leaders to collaborate on fostering diversity and inclusion, with a surprise performance by the San Francisco Gay Men’s Chorus inspiring unity and activism.
Sep 28, 2017
1,506 words in the original blog post.
PagerDuty has introduced time-based alert grouping for all standard accounts, designed to enhance incident management by reducing alert fatigue and improving triage efficiency during "alert storms." This feature automatically groups multiple alerts occurring within a specified timeframe into a single incident, thus minimizing the noise and redundancy often caused by excessive notifications. The approach allows teams to focus on resolving issues efficiently, as they receive fewer, more informative alerts that capture the entire span of the incident. By centralizing information into a single incident object, teams can streamline their response to outages, prevent redundant incidents, and maintain a clear overview of changing incident dynamics. This practical solution is likened to self-lacing sneakers, offering a simple yet effective way to address a common problem. Users have reported significant improvements in managing critical business services, with NBCNews Digital citing the prevention of thousands of redundant incidents during a beta test. With options to configure grouping periods from 2 minutes to 24 hours, this feature supports a more predictable and organized response strategy, and PagerDuty is inviting feedback and participation in a preview of future intelligent alert grouping developments.
Sep 27, 2017
823 words in the original blog post.
Configuration management tools like Chef, Puppet, and Ansible have revolutionized IT operations by enabling the automation of server provisioning and configuration, but incident management often remains a manual process. As infrastructure becomes increasingly complex with the integration of cloud servers, containers, and IoT devices, there is a growing need to automate incident management to match the scalability and efficiency of scripted infrastructure. This includes automating tasks such as alert routing, incident escalation, and alert management to reduce manual intervention and improve response times. Integrating incident management tools with ChatOps and ensuring compatibility with various monitoring systems like AWS Cloudwatch and Nagios can further streamline operations. Centralizing these tools in a cloud-based solution ensures 100% uptime, avoiding the pitfalls of relying solely on on-premises notifications. As next-generation incident management solutions become more accessible, IT administrators are encouraged to adopt these technologies to enhance productivity and effectively manage the demands of modern infrastructure.
Sep 26, 2017
857 words in the original blog post.
Enterprise IT is rapidly evolving due to technological advancements, user demands, and organizational pressures, requiring a shift from traditional centralized data processing to more agile, mobile, and focused solutions. The proliferation of mobile devices has heightened the necessity for data accessibility on-the-go, pushing organizations to adopt flexible, user-friendly systems that cater to the needs of employees who often use personal devices for work. At the organizational level, the demand for cost-effective, department-specific solutions is rising, challenging traditional enterprise software vendors to adapt. This transformation is encapsulated in the "agility-stability paradox," which requires businesses to remain responsive to technological changes and market demands while maintaining core stability and purpose. To navigate this paradox successfully, enterprises must integrate adaptive technologies and services, ensuring they can swiftly address infrastructure issues and align with both customer and employee needs, ultimately sustaining their foundational values amidst rapid changes.
Sep 19, 2017
1,031 words in the original blog post.
Inheriting a freelance development project with a disorganized codebase, poor documentation, and limited communication from previous developers can be a challenging task, as illustrated by the author's experience with a particularly troublesome project. The codebase, filled with inefficient programming patterns, required significant resources to stabilize, but the integration of incident management tools proved invaluable. These tools enabled the identification and resolution of critical issues such as database locking and memory leaks, which were otherwise difficult to detect and solve. For instance, a problematic hourly cron job causing site crashes was refactored, improving site uptime, while memory issues, including slow-loading pages due to inefficient database queries, were addressed by optimizing server configurations and caching results. The use of incident management tools not only facilitated problem-solving but also highlighted the potential for these tools to plan upgrades and scale well-architected applications effectively, underscoring their importance in both troubleshooting and project growth.
Sep 13, 2017
852 words in the original blog post.
Shadowing on-call engineers at PagerDuty revealed a work environment marked by precision, speed, and high stakes, akin to watching a medical procedure. Despite the inherent stress and disruptions to daily life, the experience was unexpectedly filled with kindness, humor, and strong support among team members. A notable instance involved a new engineer learning to resolve issues during a live incident, supported and encouraged by more experienced colleagues, illustrating the importance of humility, gratitude, and peer encouragement in creating a positive and empowering atmosphere. The culture at PagerDuty, enriched with humor and levity, helps reduce stress and fosters a collaborative team spirit, where engineers, even those not on-call, remain engaged and supportive. This environment is further enhanced by the use of gifs for humor and camaraderie, highlighting the importance of presence, encouragement, and acknowledgment in making on-call experiences more positive and manageable.
Sep 12, 2017
843 words in the original blog post.
The second annual PagerDuty Digital Operations Excellence Awards were presented at the Summit 2017 in San Francisco, recognizing outstanding customers and partners who have implemented innovative and transformative digital operations management strategies. Winners were announced in eight categories, highlighting organizations like Amazon Web Services, Atlassian, Twitter, IBM, Gap, GE Digital, Palo Alto Networks, Airbnb, and Outcome Health for their exemplary use of PagerDuty solutions. These awards celebrate achievements in areas such as innovation, digital transformation, scaling operations, customer experience, support use cases, security incident response, cultural impact, and kick-starting digital operations in startups. The awards underscore how these organizations have leveraged PagerDuty to enhance system visibility, improve operational performance, streamline incident response, and foster agile, empowering work environments.
Sep 08, 2017
1,000 words in the original blog post.
PagerDuty's second annual user conference, PagerDuty Summit, gathered DevOps practitioners and business leaders to share best practices and insights, emphasizing the importance of integrating DevOps across entire organizations beyond just engineering teams. The event highlighted key themes such as the need for all business departments, including non-technical ones, to adopt an on-call culture to enhance customer experiences. It also underscored the necessity for businesses to adapt to changing customer expectations by ensuring seamless digital operations and promoting a culture of empowerment through advanced machine learning and automation capabilities. The introduction of PagerDuty Community aims to foster collaboration and learning, while the company's commitment to the Pledge 1% initiative reflects its dedication to social responsibility. Additionally, the conference recognized innovative approaches in digital operations management through customer and partner awards, showcasing successful implementations of PagerDuty's solutions.
Sep 08, 2017
1,199 words in the original blog post.
PagerDuty announced several new product innovations at the PagerDuty Summit 2017, aimed at enhancing digital response capabilities for businesses. These innovations integrate machine learning and end-to-end response automation to address the challenges of rising infrastructure complexity and overwhelming data volumes. Key features include Alert Grouping, Similar Incidents, and a redesigned live incident page, which provide responders with intelligent, real-time decision support by organizing related alerts and offering past incident context. For response teams, new tools such as Response Plays and event routing automate incident management, allowing teams to restore service quickly and efficiently. Additionally, Dynamic Notifications enable tailored notification responses, reducing alert fatigue. The updates extend beyond technical responses, offering business-wide orchestration practices to involve all stakeholders in major incident responses, ensuring a unified focus on protecting the customer experience. These new capabilities are accompanied by updated incident response documentation and are available on various plans, with some features in customer preview and early access.
Sep 07, 2017
983 words in the original blog post.