Home / Companies / Incident.io / Blog / March 2025

March 2025 Summaries

9 posts from Incident.io

Filter
Month: Year:
Post Summaries Back to Blog
The text is a series of interviews with Lucy Jennings, an Expansion Account Manager at incident.io. The company's culture is described as transparent, kind, and collaborative, with employees encouraged to cross-collaborate and work on tasks outside their job description. Lucy shares her favorite memories, values, benefits, and experiences working at incident.io, including a team BBQ on a yacht and a high-paced day-to-day as an Expansion Account Manager. She also discusses the importance of reliable incident management tools for companies, citing the need to track time spent on tasks and make informed decisions about infrastructure investments. Lucy expresses her pride in various projects, such as the company's Status Page push, and enjoys the relaxed atmosphere of the first Friday of each month.
Mar 26, 2025 814 words in the original blog post.
We've been exploring the limitations of traditional metrics like MTTR (mean time to recovery) in measuring effective incident management. These time-based metrics are easy to calculate and provide a simple metric for comparison, but they're often misleading as they don't account for the quality of the incident response process. To address this, we built a benchmark report that focuses on quality-focused, data-backed metrics that reflect what actually happens during incident response, such as time to mobilize, time to assign a lead, frequency of updates, and out-of-hours paging. These metrics provide a more nuanced understanding of an organization's incident management process, enabling fair comparison and actionable insights for improvement.
Mar 25, 2025 737 words in the original blog post.
Integrating incident management tools with Jira enhances enterprise workflows by centralizing post-incident tasks, improving visibility, and fostering collaboration across teams. Enterprises benefit from the automated creation of action items and enhanced dashboards that provide leadership with actionable insights without system switching. Customization options in Jira allow organizations to tailor incident tracking, but challenges such as complex configurations, scaling workflows, and maintaining Jira as the "single source of truth" require careful planning. A phased integration approach, leveraging automation and templates, aids in overcoming adoption hurdles and ensures smooth implementation. This integration supports a shift from reactive incident resolution to a proactive, continuously improving process, promoting operational efficiency and accountability within large-scale enterprises.
Mar 21, 2025 988 words in the original blog post.
We're nearing the end of Q1 2025, and I'm excited about what's next for Q2 and beyond, as our product is growing rapidly with hiring ramping up. Incident management will look very different by the end of the year due to the AI explosion, so we're investing heavily in AI to make a significant difference. Our product vision focuses on the Product, which has seen 13 changelog entries, shipped over 200 fixes and features, deployed around 2700 times, and engaged with customers over 3000 times. The main areas guiding our goals are AI, On-call, and Response. We're making a huge investment in AI to make incidents easier to handle, and we've launched Scribe, which transcribes incidents, giving users peace of mind. Our next leap is Investigations, where AI completes initial investigations, allowing users to review and act. The engineering team tackles exciting challenges every day, focusing on building actionable alert insights, extending global reach, scaling alert ingestion, and adding enhancements like prompting teammates to support colleagues after a rough on-call night. We're also delivering first-class support for security incidents, creating seamless tools for curating incident timelines, developing features to build clearer post-mortem narratives, simplifying managing compliance policies, and elevating StatusPages. Our team is an energetic, tight-knit group of 80 people, maintaining incredible velocity while still finding time for dad jokes and laughs. We're proud of our culture and openly share how we operate, with a clear, ambitious vision that includes sales, Customer Success, Engineering, Product, Design, Founders working together to win.
Mar 20, 2025 1,648 words in the original blog post.
Incident debriefs can be engaging and useful by following a simple framework. The first step is to ditch the blame game and focus on understanding what happened, setting a safe space for open sharing of perspectives. Next, create a clear, factual timeline, encouraging active participation during its review. Then, focus on real-world impact, highlighting how users were affected, and evaluate your response by discussing transparency, tools, and communication. Extract actionable lessons from the incident, discuss surprises that team members feel, and commit to improvements with clear owners and due dates. Finally, finish strong by reiterating the importance of incidents as opportunities for growth, celebrating successes, and documenting and sharing insights to build a stronger team.
Mar 13, 2025 765 words in the original blog post.
Atlassian has announced that it will be shutting down Opsgenie, their popular on-call alerting tool, after June 4, 2025. This means that no new accounts will be created and the service will shut down completely by April 5, 2027. Users of Opsgenie must now plan their next steps as a key part of their incident response process is disappearing. The shutdown highlights a broader industry shift away from legacy tools, with engineering teams reconsidering their on-call management approach due to frustrations with traditional solutions such as PagerDuty and Opsgenie. These issues include rigid scheduling, clunky user interfaces, fragmented incident data, and escalating costs. In contrast, incident.io offers an integrated platform for alerts, incident coordination, and communication, providing a streamlined way to handle incidents from alert through resolution. The company is now helping users migrate away from Opsgenie to its On-Call solution, which is built for modern engineering teams and offers features such as native Slack + Teams integration, flexible scheduling, and regular improvements.
Mar 13, 2025 723 words in the original blog post.
Behind the Flame is an interview series created by incident.io, showcasing their employees' experiences and insights. The main topic of this interview is Dylan Rose Muller, a Business Development Representative at incident.io, who shares his favorite memories, values, and benefits of working for the company. He highlights the importance of teamwork, collaboration, and being part of a culture that encourages open communication and innovation. Throughout the conversation, he emphasizes the value of win together, the first Friday of the month as a benefit, and the need for reliable incident management tools to mitigate downtime and its financial impact on organizations. The interview showcases Dylan's personality, his enthusiasm for working at incident.io, and the company's mission to make a positive impact in the industry.
Mar 13, 2025 1,148 words in the original blog post.
A well-crafted runbook is a structured guide that provides clarity under pressure, reduces the risk of error, and helps Security Operations Center (SOC) teams respond faster and more confidently to incidents. A great runbook offers consistency in execution, enables faster incident resolution, accelerates onboarding for new analysts, and leaves a reliable paper trail for documentation and audits. To build an effective runbook, start with a clear purpose, break it into concise steps, use visuals to support clarity, define roles and responsibilities, add troubleshooting tips, make it a living document, and test before deployment. Effective runbooks are operational tools that strengthen a team's ability to respond, recover, and improve, and can be built using flexible tools like incident.io for free.
Mar 11, 2025 601 words in the original blog post.
PagerDuty has been the go-to on-call tool for engineering teams for years, but many teams are rethinking their use of it due to its reliability and cost issues. The company behind incident.io On-Call argues that it offers a better way forward by tackling common headaches such as rising costs, clunky service models, inflexible schedules, and a disconnect from how modern teams handle incidents. Incident.io flips this on its head by offering a full incident management platform with features like scheduling, response, and status pages in one place, providing a more streamlined experience for teams. The tool also offers better value, straightforward billing, and a single place to see what's going on, making it an attractive alternative to PagerDuty.
Mar 03, 2025 964 words in the original blog post.