Home / Companies / CloudAgent / Blog / July 2026

July 2026 Summaries

1 posts from CloudAgent

Filter
Month: Year:
Post Summaries Back to Blog
In a study examining the effectiveness of AI agents in performing cloud operations tasks, three agent harnesses—Claude Code with Opus 4.8, Codex with GPT-5.6-sol, and Cursor with Composer 2.5 Fast—were tested on ten tasks within an AWS environment. These tasks included identifying unused resources, investigating web service errors, and reviewing backup reports. Each agent executed these tasks six times, both with detailed runbooks and short prompts, resulting in a high success rate of 94% to 98%. Cursor emerged as the most efficient and cost-effective agent, while Codex excelled at following instructions accurately, and Claude provided the clearest analytical reports. The study revealed that while agents were cautious with write tasks requiring approval, they encountered issues with read-only tasks, such as silent failures during pagination. The use of detailed plans served to calibrate the agents rather than enhance their capability, suggesting that agent behavior is influenced more by temperament than raw ability. The findings indicate that AI agents are sufficiently capable and safe for daily cloud operations, though guardrails and careful verification remain necessary, particularly for read tasks.
Jul 29, 2026 1,149 words in the original blog post.