Home / Companies / Komodor / Blog / Post Details
Content Deep Dive

Building AI SRE Agents, Part 2: Leave the Laptop, Earn Trust

Blog post from Komodor

Post Details
Company
Date Published
Author
Nir Adler
Word Count
3,065
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Part two of the AI SRE agent series describes how to move a locally tested, read-only agent into a cloud runtime that can respond to real incidents while limiting operational risk. It recommends deploying the agent as a durable, event-driven service using frameworks such as Claude Agent SDK, LangGraph-based tools, AWS Strands, Microsoft Agent Framework, or Open SRE, while preserving previously validated skills, prompts, and evaluation datasets. The agent should maintain persistent incident state, receive only narrowly filtered and high-signal alerts, verify and deduplicate webhook requests, queue work to manage bursts and costs, and retain read-only access to real clusters through tightly scoped, short-lived credentials. In shadow mode, it investigates incidents and posts hypotheses and proposed fixes alongside human responders without taking action, allowing actual causes and resolutions to become ground truth for continuous evaluation. Comprehensive tracing of tool calls, decisions, latency, and token costs supports debugging, replay, and measurement of accuracy, false positives, evidence quality, and time to hypothesis. Advancement toward remediation follows an evidence-gated trust ladder from non-production shadow mode to production observation, human-reviewed proposals, and finally explicitly approved, low-risk actions on non-critical services, while broader autonomous production operation remains dependent on future enterprise controls such as governance, least-privilege access, isolation, and auditability.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.