Home / Companies / Acceldata / Blog / Post Details
Content Deep Dive

Accelerating Apache Spark with Gluten + Velox: Vectorized Execution for Big Data at Scale

Blog post from Acceldata

Post Details
Company
Date Published
Author
Senthil Kumar Balaguru
Word Count
582
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Acceldata's ODP Spark with Gluten and Velox presents a significant advancement in distributed analytics by addressing performance bottlenecks associated with Spark's traditional row-based execution model. By employing vectorized execution with columnar batches, the solution optimizes CPU cache locality and reduces function call overhead, achieving 1–3 times faster query execution and 20–30% fewer CPU cycles per row on TPC-DS 100 GB benchmarks. This approach not only enhances performance but also reduces infrastructure costs and failures due to out-of-memory errors without requiring changes to existing Spark applications. The integration of Gluten as a bridge between Spark and native engines, along with Velox's native vectorized runtime, enables seamless execution of complex analytical workloads, including aggregations, joins, and window functions. Additionally, the solution supports Apache Arrow-based zero-copy columnar data exchange and provides extensive deployment options, making it suitable for OLAP workloads with significant scalability and efficiency improvements.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 2 6,551 1,245 236 +61%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.