Home / Companies / Neo4j / Blog / Post Details
Content Deep Dive

The Spark of Neo4j

Blog post from Neo4j

Post Details
Company
Date Published
Author
Davide Fantuzzi
Word Count
945
Company Posts That Month
25
Language
English
Hacker News Points
-
Post removed?
No
Summary

The Neo4j Connector for Apache Spark is a rewritten library that leverages the new DataSource API V2, allowing for multi-language support. It was developed to provide an official library and continuous service, replacing the old connector with custom "hacky" solutions. The development process was challenging due to lack of documentation, breaking changes between minor versions, and dealing with different Scala versions supported by each Spark version. To overcome these challenges, the team created a solution that maps nodes and relationships into tables using columns for properties and IDs. They also extracted a schema from a schema-less graph, handling cases where properties may have mixed types across nodes. The connector uses the official Neo4j Java Driver and generates Cypher queries through the API. It offers examples in Scala, Python, R, and more, with plans to release support for Spark 3.0 and 3.1 soon.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.