Introducing the Neo4j Text2Cypher (2024) Dataset
Blog post from Neo4j
The Neo4j Text2Cypher (2024) Dataset is a machine learning dataset designed to help train and benchmark Text2Cypher models with ease. It was created by combining publicly available datasets, cleaning and organizing them for smoother use. The dataset consists of 44,387 instances, with 39,554 in the training split and 4,833 in the test split. The data preparation involved identifying and gathering datasets, combining and cleaning the data, creating training and test splits, and splitting remaining datasets. The resulting dataset is designed to support machine learning models that translate natural language into programming or domain-specific languages, such as turning plain text into Cypher query language.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 2,876 | 370 | 130 | -20% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.