Benchmarking Using the Neo4j Text2Cypher (2024) Dataset
Blog post from Neo4j
We explored how various fine-tuned and foundational LLM-based models perform in translating natural language questions to Cypher queries using the newly released Neo4j Text2Cypher (2024) Dataset. The results showed that closed-foundational models, such as OpenAI's GPT and Google's Gemini, demonstrated strong performance with user-friendly APIs and reliable output, though they can be costly. Previously fine-tuned models haven't quite matched these giants, but they demonstrate real potential for improvement through techniques like fine-tuning. We benchmarked four fine-tuned models and 10 foundational models to assess their performance side by side, using two evaluation procedures: translation-based evaluation and execution-based evaluation. The closed-foundational models delivered the best overall performance, with a match ratio of about 30 percent in the execution-based evaluation and outperforming previously fine-tuned models in the translation-based evaluation.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 2,876 | 370 | 130 | -20% |
| AI Model Fine-tuning | 1 | 547 | 127 | 59 | -39% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.