November 2017 Summaries
7 posts from PagerDuty
Filter
Month:
Year:
Post Summaries
Back to Blog
Achieving a positive customer service experience requires proactive management of the factors that shape it, particularly through the design of robust infrastructure and software tools. Anticipating user-interface, functional, and performance issues is crucial, as these can significantly affect customer satisfaction and loyalty. While user-interface problems may be easier to detect early through design testing, more complex functional and performance issues are harder to predict and can have long-term impacts. By employing a combination of monitoring, analytics, and rapid incident response, businesses can minimize the time between the emergence and resolution of problems, often detecting and addressing them before they become apparent to customers. Effective monitoring should capture metrics that directly or indirectly affect customer interactions, while advanced analytics can provide insights into potential performance degradations. A proactive incident management system can quickly address failures and prevent them from escalating, ensuring that issues are managed efficiently and effectively to maintain a high-quality customer experience.
Nov 29, 2017
1,141 words in the original blog post.
Cloud migration offers significant benefits such as improved agility and scalability, but it also presents challenges that organizations must address to succeed. It is crucial to transition not only systems and applications but also people and processes to embrace an agile DevOps culture that emphasizes automation, ownership, and continuous learning. Ensuring ongoing team and process health by monitoring operational metrics like MTTA and MTTR helps avoid burnout and enhance productivity. Visibility into both cloud-native and legacy systems is essential, requiring integration of monitoring tools to gain insights and enable effective incident management. A service-centric approach, which focuses on understanding service dependencies and prioritizing actionable information, can help reduce downtime and enhance customer experiences. Preparing for common pitfalls and establishing the right operational framework are key to achieving successful cloud migration and meeting business objectives.
Nov 27, 2017
1,071 words in the original blog post.
At PagerDuty, the integration of their users, customers, and buyers into a singular group provides a unique advantage, particularly as the company delves into major incident response with Network Operation Centers (NOCs). While technological advancements promise significant, yet manageable changes for NOCs, the company envisions a future where roles evolve into those like Site Reliability Engineers (SREs), who prevent and resolve glitches to maintain seamless user experiences. The transition from traditional operations to DevOps emphasizes the importance of reliability, replaceability, and routine in server management, akin to quality assurance practices. This shift is demonstrated by the example of a Los Angeles telecommunications NOC, where a structured promotion system fosters talent growth, reflecting a broader industry trend of diversifying career paths within organizations. As NOCs continue to transform, adapting to technological changes will be crucial for future success, and PagerDuty aims to support this evolution to ensure operational efficiency and resilience.
Nov 21, 2017
935 words in the original blog post.
Leveraging cloud infrastructure is increasingly seen as essential for IT organizations to enhance responsiveness to business needs, though it introduces challenges such as minimizing downtime, cutting costs, and preventing employee burnout. As organizations transition to cloud-based systems, particularly within hybrid environments combining traditional and cloud infrastructures, implementing effective incident management solutions becomes crucial to address potential issues and ensure seamless cloud migration. Migrating to the cloud offers numerous advantages, including reduced infrastructure management, increased agility, and disaster recovery capabilities, while eliminating capital investments through subscription-based services. However, the shift requires real-time visibility and effective incident response to prevent risks and ensure continuous service availability. Ensuring cloud visibility and integrating with both modern and traditional infrastructure monitoring tools is vital for managing hybrid environments, while centralized communication platforms like Slack or HipChat can support distributed workforces by integrating incident management and enhancing collaboration. Supporting teams must be equipped with the necessary tools and processes to back developers and business units, as the added complexity of cloud migration must be reliably managed to deliver business value.
Nov 14, 2017
1,170 words in the original blog post.
Achieving scalability in software deployment is crucial for business growth and requires strategic planning and execution to handle varying application demands. Legacy systems often falter under unexpected traffic spikes, leading to customer loss and operational challenges, while even cloud-based applications can suffer from bottlenecks if not properly designed. Scalability can be attained through the implementation of the right tools, processes, and team structures, including the use of DevOps practices, Infrastructure-as-Code, and cloud container platforms. Effective scalability also depends on a cohesive team structure involving developers, IT operations teams, and site reliability engineers (SREs), who work together to ensure reliability and efficiency. The transition from traditional methods to modern ITOps and reliability engineering necessitates improved communication and collaboration, alongside the adoption of practices like canary releases to optimize performance and minimize disruptions. Embracing these changes enhances productivity and enables immediate value realization, akin to the benefits seen in agile development.
Nov 09, 2017
951 words in the original blog post.
PagerDuty is eagerly preparing for AWS re:Invent 2017, where they will engage with the Amazon Web Services community to explore ways to enhance developer productivity, network security, and application performance. Positioned at Booth #2628 in the Venetian, PagerDuty will showcase its platform's integrations with technologies such as AWS CloudWatch, AppDynamics, and Datadog, emphasizing how it facilitates seamless cloud migration and empowers teams with a DevOps culture. Attendees can experience demos of PagerDuty's machine learning and response automation innovations and learn about its Digital Operations Management platform, which combines cloud and on-premises data with incident response best practices. The company is also hosting several events, including a pub crawl co-hosted with Datadog and Chef at AquaKnox, where attendees can network and enjoy cocktails. Additionally, Mark Gabbard from PagerDuty will discuss four strategies for cloud transformation using insights from AWS CloudWatch and other tools during a session at the Kumo Theater.
Nov 07, 2017
550 words in the original blog post.
In the modern business landscape, organizations are under immense pressure to meet high consumer expectations and manage vast amounts of data, with brand value heavily influenced by the quality of customer experiences. As software becomes a primary medium for interaction, the ability to innovate swiftly and maintain high-quality digital services is crucial for survival in a competitive environment where failure to utilize data effectively can lead to significant losses and reputational damage. Digital Operations Management, which integrates machine learning, automation, and DevOps-centric workflows, plays a pivotal role in enabling teams to respond promptly to critical issues, thereby improving business outcomes. Companies like IBM, Airbnb, and General Electric are among the over 9,500 customers using PagerDuty to enhance their digital operations. A free trial offers the opportunity to experience how data can be transformed into real-time insights and actions, with additional resources available on their Digital Operations Management page.
Nov 06, 2017
239 words in the original blog post.