Home / Companies / Incident.io / Blog / July 2025

July 2025 Summaries

6 posts from Incident.io

Filter
Month: Year:
Post Summaries Back to Blog
James Jarvis reflects on his decision to join incident.io as a product engineer, driven by a desire to be part of the AI revolution and inspired by the team's impressive presence and productivity. He highlights the rewarding nature of his role, which allows for immediate impact and direct feedback, contrasting it with experiences at other companies. Jarvis notes the fierce competition in the AI-driven incident management space, underscoring incident.io’s commitment to innovation and rapid development. The company's strong values, such as trust and collaboration, create a supportive environment where new employees are quickly integrated and given significant responsibilities. Jarvis appreciates the opportunity to work with talented colleagues who contribute to building an industry-leading product, while also acknowledging the occasional imposter syndrome that comes with being part of such a capable team.
Jul 31, 2025 1,472 words in the original blog post.
At incident.io, an "office-first" mindset is embraced, highlighting the benefits of in-person work such as faster collaboration, effortless learning, and stronger connections among team members. The office environment is seen as fun and energizing, facilitating quick problem-solving through spontaneous interactions and shared insights. While remote work remains an option, with flexibility allowing employees to tailor their schedules around personal needs, many prefer the office for its ability to foster trust, accountability, and camaraderie. The positive atmosphere, characterized by a mix of work and lighthearted moments, enhances productivity and enjoyment, making the office a preferred setting for most employees despite the availability of remote work options.
Jul 24, 2025 904 words in the original blog post.
At incident.io, the goal to achieve a five-minute deploy time became increasingly challenging as the codebase and test suite expanded, initially resulting in deployment times exceeding 11 minutes. To address these delays, the team transitioned their CI/CD pipeline to Buildkite, allowing them to run their own high-powered servers and implement effective caching strategies for both Go and TypeScript codebases. This shift significantly reduced test times, with production deploys decreasing by 27% and hotfix deploys achieving the target of under five minutes. However, the transition faced challenges such as managing server resources, optimizing disk I/O, addressing flaky tests, and resolving cache issues. Despite these obstacles, the improvements align with the company’s core value of increasing pace and enable faster delivery of new features to customers, though the quest for consistently achieving the five-minute deploy target for regular releases continues.
Jul 22, 2025 1,900 words in the original blog post.
At incident.io, the on-call system is designed to ensure continuous product reliability and improve both internal processes and the product itself, with all engineers participating in a rotation. The schedule begins on Friday evening and runs for a week, with the rationale that starting on the weekend reduces the week's looming pressure. Support systems are in place, allowing team members to request cover or escalate issues if necessary, and new engineers are eased into the rotation with mentorship. Compensation for on-call duties is included in overall pay, with additional hourly rates for time outside regular hours, and team members are encouraged to rest after challenging nights. The company fosters a supportive environment by emphasizing mutual coverage, trusting engineers to manage their readiness and responsibilities while maintaining a healthy work-life balance.
Jul 18, 2025 908 words in the original blog post.
In 2025, a curated list of top engineering voices highlights influential figures in the tech industry who share their insights and experiences to enhance the community. These engineers, including Camille Fournier, Cassidy Williams, and Charity Majors, contribute through blogs, podcasts, and public talks, covering topics such as technical leadership, balancing productivity with mental health, platform engineering, and system reliability. The list features diverse perspectives from industry leaders like Dr. Claire Knight, Jaana Dogan, and Julia Evans, who offer practical advice and demystify complex technical concepts. This collection of voices emphasizes the importance of thoughtful, inclusive leadership and sustainable work cultures while making deep technical knowledge accessible and engaging. The individuals are celebrated for their ability to inspire, educate, and foster a more connected and thoughtful engineering community.
Jul 18, 2025 3,145 words in the original blog post.
Incident.io has launched AI SRE, an AI-driven Site Reliability Engineering tool designed to streamline incident management by working alongside human engineers to quickly identify, diagnose, and resolve system issues. This new tool, developed over four years, aims to reduce downtime and alert fatigue by autonomously handling incidents and only involving engineers when human intervention is necessary. AI SRE operates within platforms like Slack and Microsoft Teams, integrating data from various sources to offer transparent and traceable solutions, thereby allowing engineers to focus on strategic, roadmap-related tasks rather than being bogged down by routine incidents. The platform promises faster incident resolution and improved efficiency by instantly investigating issues, providing root cause analyses, and suggesting actionable fixes, drawing positive feedback from early adopters. This development follows the company's recent $62 million Series B funding, emphasizing its commitment to revolutionizing incident management with AI at its core.
Jul 02, 2025 1,305 words in the original blog post.