October 2026 Summaries
4 posts from Komodor
Filter
Month:
Year:
Post Summaries
Back to Blog
Autonomous AI SRE agents require operational safeguards beyond effective prompting when they begin responding directly to production alerts. The recommended architecture uses an intake layer to normalize, deduplicate, group, and route alerts, plus queues, concurrency controls, retries, and strict time, tool-call, and token budgets to prevent alert storms and runaway costs. Agents should be deployed as version-pinned, observable workloads with least-privilege access, network restrictions, cloud-hosted model endpoints, credential-free prompts, and isolated ephemeral sandboxes for executing untrusted code or testing fixes. Comprehensive traces should record incidents, model and tool actions, approvals, and outcomes for auditing and change-management requirements. Continuous improvement depends on evidence-backed, aging incident memory, engineer feedback at resolution, and evaluation sets built from confirmed incidents, while any prompt, model, or skill change is tested and shadowed before promotion. Autonomous remediation should expand gradually by action class according to demonstrated accuracy, limited blast radius, tested rollback procedures, and approval tiers, with humans retaining oversight for higher-risk changes.
Oct 08, 2026
2,418 words in the original blog post.
Context engineering for AI agents involves selecting the smallest, most relevant set of instructions, tools, knowledge, memories, live data, and handoffs that an agent needs at each step, rather than simply expanding its context window. In operational settings, the article argues that context must function as shared infrastructure so agents can build on reviewed findings from previous incidents instead of repeatedly starting from scratch. It distinguishes stable, always-loaded material such as instructions and safety rules from information retrieved on demand, including runbooks, logs, metrics, and specialized procedures. Excessive, irrelevant, or outdated context can reduce reliability, increase hallucinations, and raise token, latency, and tool-use costs, even when an agent reaches the correct conclusion. Recommended practices include filtering raw data before presenting it to models, narrowing retrieval by metadata, using specialist agents with shared findings, storing concise facts rather than transcripts, reviewing and expiring memories, and measuring context costs per tool call. Komodor describes its platform as implementing these ideas through a shared knowledge graph, curated knowledge base, agent memory, change history, and review processes intended to keep operational knowledge current, permissioned, and reusable across agents.
Oct 07, 2026
2,084 words in the original blog post.
Komodor CEO Ben Ofiri describes the company’s Agentic Operations Platform as an extension of its AI SRE work, designed to help medium and large enterprises build, deploy, govern, evaluate, and optimize AI agents for operational tasks such as incident response, troubleshooting, and cost management. He argues that while enterprises increasingly seek agentic automation for DevOps and SRE, moving agents from prototypes into production introduces major challenges involving governance, visibility, reliability, performance measurement, and token-related costs. Komodor positions its platform between closed, vendor-provided AI SRE products that may be difficult to customize and generic agent frameworks that require organizations to develop integrations, workflows, evaluation methods, and operational expertise themselves. The platform offers prebuilt SRE agents alongside tools for importing or creating custom agents, selecting deployment locations, enforcing policies, running evaluations, and optimizing costs. Ofiri attributes the approach to a combination of AI expertise and practical SRE knowledge, emphasizing that regulated organizations often require auditability and human oversight. He predicts that over the next two to three years, AI will automate much of operational work, particularly incident detection, investigation, and triage, while some remediation will continue to require human involvement.
Oct 06, 2026
2,687 words in the original blog post.
Komodor migrated its metrics platform from Amazon Timestream to ClickHouse Cloud after AWS restricted new access to Timestream LiveAnalytics and the company encountered high, unpredictable costs, limited data mutability, aggregation constraints, weak observability, and no local testing environment. The seven-month project moved approximately 5.9 billion metric records per day through a feature-flagged, reversible process involving dual writes, shadow queries, parity validation, phased account cutovers, and eventual Timestream decommissioning, allowing production traffic to continue without customer-facing downtime. After testing InfluxDB 3, TimescaleDB, and ClickHouse under production-scale traffic, the team selected ClickHouse for its ingestion capacity, SQL support, low latency, compression, materialized views, Docker-based local development, and projected monthly cost of $5,000–$6,000 versus about $38,000 for Timestream. Extensive unit, component, parity, and live shadow testing, supplemented by AI-assisted SQL generation and validation reporting, revealed unexpected data discrepancies caused by differing duplicate-record handling between databases and flaws in a custom metrics collection plugin. Across 10.3 million production query comparisons, ClickHouse reduced average query latency from 906 milliseconds to 33 milliseconds and p95 latency from 1,274 milliseconds to 116 milliseconds, while lowering database costs by roughly 85 percent and enabling schema changes, metadata joins, and broader future data-processing capabilities.
Oct 01, 2026
3,850 words in the original blog post.