Eric Schwartz on what it takes to run an AI SRE at petabyte scale
Blog post from WorkOS
Traversal, a Series A company serving large enterprises, develops an AI site reliability engineering platform designed to automate incident diagnosis, alert triage, and eventually remediation as AI coding tools increase software output without reducing operational troubleshooting work. Its approach relies on continuously integrating, compressing, and indexing customers’ massive telemetry volumes before an agent investigates, rather than querying billions of logs live, while combining specialized SRE workflows, tools, and prompts with different frontier and open-source models according to incident severity and cost. The company says deployments can generally reach production within a week, though forward-deployed engineers still support optimization, and it positions its focus on large-scale data ingestion as a key distinction from coding agents that can also perform debugging tasks. Customers control a staged autonomy model, ranging from diagnosis only to agent-created fixes and pull requests requiring human review, with alert-noise reduction often automated earlier because teams may not actively monitor high-volume channels. Traversal has also shifted from fully on-premises deployments to standard SaaS and AWS-based bring-your-own-cloud options, reflecting enterprise requirements around security, infrastructure, and trust.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.