Home / Companies / PagerDuty / Blog / July 2017

July 2017 Summaries

9 posts from PagerDuty

Filter
Month: Year:
Post Summaries Back to Blog
PagerDuty has announced its expansion into the Asia-Pacific region with a new office in Sydney, Australia, highlighting the strategic importance of this market for its long-term growth. This move is complemented by the release of the State of Digital Operations in Australia Report, which surveyed over 200 IT personnel and 300 consumers, revealing a disconnect between consumer expectations and the ability of IT organizations to deliver seamless digital services. The report underscores the need for enhanced digital operations management, noting that resolving consumer-impacting incidents takes IT teams significantly longer than consumers are willing to wait, potentially leading to customer dissatisfaction and revenue loss. Despite a majority of IT personnel believing their organizations are equipped to support digital services, many still face frequent customer-impacting incidents. The findings emphasize the importance of IT operations in ensuring seamless service delivery, which is crucial for maintaining customer loyalty and driving business revenue. As digital consumerism increases, organizations must improve operations and productivity to meet rising expectations, with PagerDuty positioned as a leader in helping companies quickly resolve business-impacting incidents to provide a flawless customer experience.
Jul 31, 2017 743 words in the original blog post.
Joining PagerDuty as an Engineering Manager during its rapid growth phase offered a unique opportunity to explore the structural challenges of scaling an organization. Initially, the engineering team faced significant siloing issues between various departments and locations, leading to inefficiencies and communication gaps. Efforts to address these issues included experimenting with different organizational structures, such as a matrix organization, which introduced dual reporting and further complications. By 2016, the company transitioned to an Agile organization with vertical product Scrum teams and horizontal teams for cross-cutting concerns, adopting a true DevOps approach where teams owned their code and infrastructure. This shift allowed for greater team autonomy, faster iteration, and minimized dependencies, leading to improved productivity and collaboration. The company continued to refine its approach, embracing a tribal model inspired by Spotify to enhance shared ownership and vision, while focusing on continuous learning and adaptability to maintain agility and innovation in a high-growth environment.
Jul 27, 2017 2,508 words in the original blog post.
PagerDuty Summit '17, scheduled to take place in San Francisco from September 5-7 at Pier 27, aims to revolutionize operations by providing attendees with insights and tools to enhance their business responses to various issues. The event is designed for professionals across development, IT operations, customer support, and executive roles, offering two main tracks: DevOps Best Practices and Leadership in a Digital World. Attendees can gain practical knowledge from PagerDuty University courses on incident response and improving on-call life, while the Breakathon provides a hands-on opportunity to diagnose and fix service failures in a competitive setting. The summit promises a wealth of content, networking opportunities, and a chance to engage with the latest in operational strategies, making it a must-attend for those looking to innovate and deliver exceptional customer experiences.
Jul 20, 2017 707 words in the original blog post.
Continuous integration aims to automate builds and tests to enhance efficiency and quality in development pipelines, yet challenges can arise with the increased pace of updates. Integrating incident management from the outset in the continuous integration process can enhance accountability, visibility, and transparency, enabling teams to handle problems efficiently when they occur. This integration ensures that both Development and Operations teams collaborate effectively, fostering shared responsibility for uptime and allowing quick troubleshooting without blame. Incident management not only enforces a culture of quality and preparedness but also necessitates transparent communication and unified monitoring metrics across teams to ensure everyone is informed and ready to respond to issues. Having a centralized hub for monitoring data and a defined process for issue resolution is essential to making data actionable and ensuring swift responses to potential problems. For a smooth DevOps transformation, continuous integration and incident management must work together, providing significant benefits to team dynamics and application reliability.
Jul 18, 2017 786 words in the original blog post.
Introducing DevOps practices into an organization requires demonstrating their value, especially since change can often be met with resistance. The process involves understanding current workflows, identifying repetitive tasks that can be automated, and promoting a cultural shift in communication and collaboration. By focusing on small, manageable changes, such as automating routine tasks or streamlining onboarding processes, teams can gradually embrace DevOps principles. Sharing the successes and efficiencies gained with other teams can help build momentum for wider adoption. Incremental improvements and cultural integration of DevOps practices can save time and resources, ultimately enhancing productivity and innovation.
Jul 13, 2017 1,041 words in the original blog post.
PagerDuty's "Failure Fridays" initiative, started in 2013, involves weekly controlled fault injections into their production environment to test and improve system resilience without affecting customers. Over the years, this practice has evolved, incorporating elements of Chaos Engineering and expanding from single service tests to infrastructure-wide simulations, including Availability Zone (AZ) and Region failures. This method has not only enabled PagerDuty to identify and rectify potential issues before they impact users but has also fostered internal trust and operational improvements. The initiative has led to significant process automation and documentation, including the development of tools like "Reboot Roulette" and "Chaos Cat" for fault injection. By June 2017, PagerDuty had conducted 121 sessions, injected 644 faults, and created over 200 tickets to address identified issues, illustrating how such stress tests can enhance both software delivery and team cohesion.
Jul 12, 2017 730 words in the original blog post.
A release in product development represents a set of customer-visible and operational features that deliver a new or improved product capability, showcasing a commitment to enhancing customer experience and interaction. Releases can be either date-driven, influenced by marketing factors, or scope-driven, focusing on specific features valuable to users. Agile methodologies treat scope as a variable, allowing for flexibility in planning and prioritization, with release forecasting playing a crucial role in coordinating launches, managing dependencies, and aligning team efforts. This practice supports cross-team collaboration and decision-making, ensuring stakeholders are informed and expectations are managed effectively. Release forecasting is not about exact precision but rather providing a framework for planning and aligning team goals, fostering buy-in, and enhancing the accuracy and empowerment of those involved.
Jul 11, 2017 796 words in the original blog post.
The rapid expansion of the threat landscape, characterized by frequent and potent vulnerabilities such as ransomware attacks, challenges ITOps teams to manage an increasing load of servers, applications, and endpoints while maintaining security. As organizations adopt agile ITOps methodologies, integrating containers and public cloud resources presents new security challenges that require a multifaceted SecOps strategy for full stack visibility and effective incident resolution. This involves simplifying SecOps stacks to reduce alert noise and enhance actionability, incorporating tools to prevent and manage threats like crypto-ransomware, and establishing a central incident management solution for enriched alerts and streamlined remediation processes. By leveraging syslog configurations, SNMP traps, and third-party intrusion analysis systems, organizations can enhance threat intelligence and reduce alert fatigue. Additionally, they must adapt these strategies for hybrid or public cloud environments using tools like Azure Alerts and AWS Cloud Watch. Ultimately, maintaining simplicity, visibility, noise reduction, and actionability is essential for effective security incident response, as outlined in resources such as PagerDuty’s open-source documentation.
Jul 07, 2017 1,103 words in the original blog post.
Modern IT infrastructure has evolved significantly from the hardware-focused systems of the past, transitioning into a predominantly software-driven environment where the line between hardware and software is increasingly blurred. This shift is characterized by the stabilization of hardware advancements, like processor speed and RAM, while software undergoes rapid changes, leading to innovations such as software-defined networking, virtualization, and cloud computing. These developments have minimized the hardware-imposed lag on infrastructure, allowing for more dynamic and flexible operations. As a result, today's IT infrastructure is largely virtualized, utilizing continuous delivery and event-driven automation to manage development and deployment processes efficiently. This virtual environment frees computing from traditional hardware constraints, paving the way for potential future advancements in virtual reality and automation that could transform both technology and human experiences.
Jul 06, 2017 1,252 words in the original blog post.