Building AI SRE Agents, Part 1: Start Local, Break Things, Learn Fast
Blog post from Komodor
AI-powered Site Reliability Engineering (SRE) agents are designed to handle the overwhelming signals generated by modern infrastructure, providing first-response incident management by triaging alerts, correlating logs and metrics, and offering remediation suggestions without replacing human engineers. The initial stage of developing these agents involves local experimentation, allowing teams to understand and trust the system in a non-production environment where mistakes are low-cost. This process involves using tools like Claude Code to build agent behavior, create skills for automating routine tasks, and refine performance against a set of known test incidents to ensure reliability before moving to production-level deployments. The goal is to establish a propose-don’t-execute methodology, where the AI suggests actions that are reviewed by humans, helping to prevent premature deployment into sensitive environments. This careful approach, starting locally and gradually moving to cloud environments, is critical for avoiding common pitfalls such as over-reliance on agent-led workflows that do not meet return on investment (ROI) expectations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MCP | 4 | 3,533 | 369 | 145 | -53% |
| Kubernetes | 2 | 1,260 | 165 | 75 | -41% |
| Secrets Management | 2 | 1,384 | 221 | 91 | -44% |
| AI Model Fine-tuning | 1 | 402 | 99 | 46 | -46% |
| Harness engineering | 1 | 137 | 67 | 36 | -46% |
| LLM | 1 | 3,751 | 612 | 168 | -39% |
| Observability | 1 | 1,844 | 344 | 128 | -56% |
| Vector Search | 1 | 1,111 | 224 | 91 | -41% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.