Home / Companies / ScyllaDB / Blog / Post Details
Content Deep Dive

Apache Spark at ScyllaDB Summit, Part 2: Tips for Building Resilient Pipelines

Blog post from ScyllaDB

Post Details
Company
Date Published
Author
Peter Corless
Word Count
1,850
Company Posts That Month
14
Language
English
Hacker News Points
-
Post removed?
No
Summary

At the ScyllaDB Summit 2018, Google’s Holden Karau presented on building resilient pipelines in Apache Spark, addressing the challenges of Spark failures and outlining strategies for recovery. As an open-source advocate and Spark committer, Karau highlighted the complexities and redundancies within Spark’s components, such as multiple machine learning and streaming engines, which can complicate pipeline recovery. She demonstrated the construction of a recoverable Wordcount example to illustrate the process, emphasizing the importance of checkpoints and handling success markers to mitigate failure impacts. Despite Spark's lazy evaluation leading to late error detection, Karau proposed techniques like caching and asynchronous saving to improve resilience, while also acknowledging the limitations and potential inefficiencies of these solutions. She stressed the importance of using job IDs to manage separate backfills without data interference and concluded by underscoring the necessity of testing and focusing recovery efforts on critical parts of the pipeline.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 3 531 163 60 +5%
Data Pipeline 1 47 20 12 -65%
Kubernetes 1 501 76 32 -37%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.