Home / Companies / Neo4j / Blog / Post Details
Content Deep Dive

Introducing the Neo4j Text2Cypher (2024) Dataset

Blog post from Neo4j

Post Details
Company
Date Published
Author
Makbule Gulcin Ozsoy
Word Count
732
Company Posts That Month
16
Language
English
Hacker News Points
-
Post removed?
No
Summary

The Neo4j Text2Cypher (2024) Dataset is a machine learning dataset designed to help train and benchmark Text2Cypher models with ease. It was created by combining publicly available datasets, cleaning and organizing them for smoother use. The dataset consists of 44,387 instances, with 39,554 in the training split and 4,833 in the test split. The data preparation involved identifying and gathering datasets, combining and cleaning the data, creating training and test splits, and splitting remaining datasets. The resulting dataset is designed to support machine learning models that translate natural language into programming or domain-specific languages, such as turning plain text into Cypher query language.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 2,876 370 130 -20%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.