Apache Spark at ScyllaDB Summit, Part 2: Tips for Building Resilient Pipelines
Blog post from ScyllaDB
At the ScyllaDB Summit 2018, Google’s Holden Karau presented on building resilient pipelines in Apache Spark, addressing the challenges of Spark failures and outlining strategies for recovery. As an open-source advocate and Spark committer, Karau highlighted the complexities and redundancies within Spark’s components, such as multiple machine learning and streaming engines, which can complicate pipeline recovery. She demonstrated the construction of a recoverable Wordcount example to illustrate the process, emphasizing the importance of checkpoints and handling success markers to mitigate failure impacts. Despite Spark's lazy evaluation leading to late error detection, Karau proposed techniques like caching and asynchronous saving to improve resilience, while also acknowledging the limitations and potential inefficiencies of these solutions. She stressed the importance of using job IDs to manage separate backfills without data interference and concluded by underscoring the necessity of testing and focusing recovery efforts on critical parts of the pipeline.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 3 | 531 | 163 | 60 | +5% |
| Data Pipeline | 1 | 47 | 20 | 12 | -65% |
| Kubernetes | 1 | 501 | 76 | 32 | -37% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.