Home / Companies / Komodor / Blog / August 2026

August 2026 Summaries

2 posts from Komodor

Filter
Month: Year:
Post Summaries Back to Blog
Komodor will host an online webinar on September 23, 2026, led by CTO and co-founder Itiel Shwartz, to introduce its Agentic Operations Platform and discuss the future of agentic AI in production operations. The session will cover the company’s experience managing incidents in large production environments, the limitations and opportunities of current agentic AI solutions, and a live platform demonstration. Attendees will see how to use customizable solution modules and more than 50 prebuilt agents, configure remediation workflows for autonomous execution or human approval, and build or import custom agents with controlled access to models, tools, skills, and permissions. The webinar will also demonstrate fleet management, continuous LLM-as-judge evaluations, shadow experiments for comparing workflow changes before deployment, and organization-wide governance controls for model use, spending, and actions requiring approval.
Aug 25, 2026 481 words in the original blog post.
Part two of the AI SRE agent series describes how to move a locally tested, read-only agent into a cloud runtime that can respond to real incidents while limiting operational risk. It recommends deploying the agent as a durable, event-driven service using frameworks such as Claude Agent SDK, LangGraph-based tools, AWS Strands, Microsoft Agent Framework, or Open SRE, while preserving previously validated skills, prompts, and evaluation datasets. The agent should maintain persistent incident state, receive only narrowly filtered and high-signal alerts, verify and deduplicate webhook requests, queue work to manage bursts and costs, and retain read-only access to real clusters through tightly scoped, short-lived credentials. In shadow mode, it investigates incidents and posts hypotheses and proposed fixes alongside human responders without taking action, allowing actual causes and resolutions to become ground truth for continuous evaluation. Comprehensive tracing of tool calls, decisions, latency, and token costs supports debugging, replay, and measurement of accuracy, false positives, evidence quality, and time to hypothesis. Advancement toward remediation follows an evidence-gated trust ladder from non-production shadow mode to production observation, human-reviewed proposals, and finally explicitly approved, low-risk actions on non-critical services, while broader autonomous production operation remains dependent on future enterprise controls such as governance, least-privilege access, isolation, and auditability.
Aug 20, 2026 3,065 words in the original blog post.