Why OOMKilled Spark Executors on Kubernetes Are Harder to Diagnose Than They Should Be
Blog post from Acceldata
Spark executors running on Kubernetes may be terminated by OOMKilled events when containers exceed their memory limits, a process managed by the Kubernetes layer rather than Spark itself. This creates a diagnostic challenge as these events are logged outside of Spark’s native visibility tools, such as the Spark UI and History Server, which only show executor loss without the root cause. The evidence for these terminations exists in Kubernetes-specific logs and events, requiring engineers to correlate data across multiple tools to identify the cause. A comprehensive observability stack is necessary to bridge this gap, integrating signals from both Spark and Kubernetes to provide a unified view of infrastructure health and resource usage, thereby facilitating quicker root cause analysis and preventing recurrence. Tools like xLake can offer this integration, reducing the manual context switching and investigation time needed to diagnose OOMKilled events.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 32 | 1,965 | 371 | 106 | -15% |
| Observability | 1 | 3,421 | 707 | 180 | -24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.