Home / Companies / ClickHouse / Blog / Post Details
Content Deep Dive

Can LLMs replace on call SREs today?

Blog post from ClickHouse

Post Details
Company
Date Published
Author
Lionel Palacin and Al Brown
Word Count
14,848
Company Posts That Month
24
Language
English
Hacker News Points
-
Post removed?
No
Summary

The text outlines an experiment conducted by ClickHouse to evaluate whether AI-powered observability, specifically through large language models (LLMs), can effectively replace Site Reliability Engineers (SREs) in performing root cause analysis (RCA). Despite the potential of LLMs like Claude Sonnet 4, OpenAI's GPT models, and Gemini 2.5 Pro, the study concluded that these models are not yet capable of autonomously identifying root causes in complex, real-world scenarios without guidance, even though they can assist with documentation tasks such as drafting RCA reports. The experiment revealed that while models could sometimes pinpoint issues, their performance was inconsistent, largely due to a lack of context and domain specialization. Furthermore, the unpredictability in token usage and cost presents challenges for integrating LLMs into automated observability workflows. The study suggests that the current best approach is a collaborative one that combines human engineers with fast, scalable observability tools and LLMs for supportive tasks, allowing for more efficient and accurate incident resolution.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 79 3,922 600 189 -6%
Observability 50 1,883 347 119 -9%
OpenTelemetry 30 403 61 26 -39%
MCP 12 3,840 275 112 +19%
Kubernetes 10 986 177 85 -38%
Real-time 3 4,334 965 217 -7%
AI Model Fine-tuning 2 568 107 59 -14%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.