May 2024 Summaries
7 posts from Incident.io
Filter
Month:
Year:
Post Summaries
Back to Blog
In this episode of Future, Toby Jackson, Global SRE Team Lead at Future, shares his insights on why taking a cookie-cutter approach to incident management is not advisable as it fails to account for the unique context and requirements of each specific situation.
May 29, 2024
127 words in the original blog post.
Skyscanner, a leading travel search engine, was looking to improve its scalability and efficiency in managing its large volumes of data. The company adopted incident.io, a cloud-based platform that enables real-time data processing and analytics. By adopting incident.io, Skyscanner aimed to enhance its ability to handle complex data sets and provide better insights to users. The adoption process involved an evaluation of the pros and cons of buying versus building a custom solution, ultimately leading to the decision to buy incident.io. The conversation between Chris Evans, CPO of incident.io, and John Paris, Principal Engineer at Skyscanner, provides valuable insights into the challenges faced by the company and how they overcame them through the adoption of incident.io.
May 21, 2024
170 words in the original blog post.
In this episode of The Debrief, we discuss the importance of staying true to your product principles when building AI features, as companies grapple with incorporating AI into their products for customer benefit. Building AI features isn't a decision made on a whim, but rather requires careful consideration of how to do it right. Companies should weigh the decision to build with AI against their existing product principles, which can impact the overall customer experience and success of the feature.
May 14, 2024
188 words in the original blog post.
When an incident strikes, the importance of psychological safety becomes starkly apparent as teams face mounting pressure and stress, which can intensify if individuals are afraid to speak up or make mistakes without fear of repercussions. Psychological safety is critical for effective incident management, promoting a culture of deep learning and blamelessness, where responders feel confident in their approach to problem-solving. A lack of psychological safety can lead to challenges across the entire incident lifecycle, including being afraid to report issues, feeling anxious about going on call, or sidestepping important work. To promote psychological safety, it's essential to create a culture where individuals feel seen and supported, making it clear that it's okay to not have the answers, looking at systems rather than individuals, facilitating open communication, encouraging asking questions, and removing undue burdens on team members by allowing them to ask for help without repercussions. By establishing this kind of culture, organizations can truly thrive and enable their teams to do their best work.
May 13, 2024
1,424 words in the original blog post.
We experimented with various AI-powered features at incident.io, learning that investing in team expertise and tooling upfront is crucial for successful implementation. We found it helpful to discover what's technically possible through short experiments and iterated quickly to build conviction in the team about where AI-powered features would be most compelling. Investing in tools and developer experience up front allowed us to experiment faster and test more ideas, resulting in a significant reduction in development time. It's equally important to invest in tooling as it is to invest in the team. We also learned that having principled product principles, such as keeping a human in the loop and subtly automating existing workflows, can help narrow down what to build and how to build it. Additionally, using large language models during development can be incredibly powerful, but it's essential to stay up-to-date with foundation model changes and focus on the big picture. Launching AI-enabled features in phases, being close to customers, and not waiting for perfection are also critical for success. Ultimately, building AI features requires a combination of technical expertise, product principles, and a willingness to learn and adapt.
May 06, 2024
1,229 words in the original blog post.
In this episode of The Debrief, SRE Dan Slimmon discusses his "clinical troubleshooting" framework for responding to incidents, highlighting its benefits and the importance of effective collaboration in resolving complex issues.
May 06, 2024
185 words in the original blog post.
The on-call experience can be stressful for engineers, but with the right approach, it can be manageable and empathetic. Simplifying and streamlining the process is key, including embracing blamelessness, which allows engineers to feel empowered to improve the process themselves. Communication is also crucial, involving well-defined roles and avoiding finger-pointing. The user experience should not get in the way of incident management tools, and good UX can make a significant difference. Finally, it's essential to consider the impact on personal life, offering flexible schedules and allowing engineers to excuse themselves when needed to reduce stress.
May 02, 2024
1,007 words in the original blog post.